Stateless MCP, 80% price cuts, and 37 agent projects in one day. Tech & AI Weekly by Eli, July 31 2026
Hi! The LLM price war arrived this week. OpenAI made its first major cut to GPT-5.6 prices — Luna dropped 80%. Chinese model competition gets the credit.
AI & GenAI
GPT-5.6 Luna drops 80% in price — and DeepSeek lit the fire. OpenAI used GPT-5.6 Sol to optimize its own inference code and GPU kernels, cutting serving costs by 20% and passing the savings on. The timing is not a coincidence. If you’ve been eyeing GPT-5.6 but watching the bill, check the new numbers. Read more
Claude Opus 5 launched this week. Near Fable 5 intelligence at half the cost, with new benchmarks in coding and scientific reasoning. It’s also on Bedrock if that’s your stack. I’ve been testing it on complex agent chains — the planning quality is noticeably different. Read more
AWS AgentCore Gateway now supports the MCP 2026-07-28 spec. This is the biggest MCP revision yet: the protocol is now stateless and runs on standard HTTP infrastructure, no centralized hub required. If you’re on AgentCore, one UpdateGateway call gets you there. Read more
Diagrid lets failed agents resume mid-task. Restarting from zero when an agent fails halfway through a complex workflow is one of the more painful production problems. This adds resume-from-failure infrastructure. Worth reading if you’re running long-running agents in production. Read more
Cloud & DevOps
GitHub Stacked PRs are now in public preview. Break large changes into smaller, ordered PRs that teams can review in parallel and merge in one click. Existing branch protections apply automatically. If you’re shipping AI-generated code in batches, this is a meaningful workflow upgrade. Read more
AWS Lambda removes the per-region aggregate storage quota. You can now reference deployment packages directly from your own S3 bucket, and the default managed storage went from 75 GB to 300 GB. Individual function size limits are unchanged — but if you’ve been hitting the account-level ceiling, this clears it. Read more
Data & ML
Beyond RAG: task-aware knowledge compression (TAKC) on AWS. Standard RAG retrieves chunks and hopes the model connects the dots across documents. TAKC pre-compresses the knowledge base into task-specific representations at 8x–64x compression, routed by query complexity. If analytical reasoning across documents is your bottleneck, the vanilla RAG fix isn’t more vectors — it’s this. Read more
From me this week
Hack the Video Agent Context Graph. Yesterday at AWS Builder Loft in San Francisco I co-organized this hackathon and it was incredible. 151 participants showed up and built 37 projects in one day, every single one using Strands Agents, OpenAI, Neo4j, and TwelveLabs together. Four developer advocates from four companies made it happen. Huge thanks to Adam Chan and HackerSquad for organizing, to judges Mike Chambers and Asako Hayase, and to all the sponsors who brought expertise and credits. See the recap
I hope you enjoy this as much as I enjoy putting it together every week. It’s my little weekly de-stress moment. Let’s stay informed together and learn something new along the way.
See you next Friday,
Eli
Dev.to · LinkedIn · GitHub · Twitter/X · Instagram · YouTube