Elizabeth Fuentes Leone
Elizabeth Fuentes Leone Elizabeth Fuentes Leone

Tutorials, deep dives, and hands-on guides on AI Agents, RAG systems, and cloud development by Elizabeth Fuentes Leone.

🏠 Home 📝 Blog 🏷️ Tags 👩‍💻 About 🎤 Event Resources 📨 Subscribe
AI agents now control lab equipment. Tech & AI Weekly by Eli, August 29 2026

AI agents now control lab equipment. Tech & AI Weekly by Eli, August 29 2026

Aug 29, 2026
newslettertechai

[Read on the web]({{ BLOG_URL }})

Hi—

This week Claude learned to control lab equipment. Not a chatbot improvement, not a benchmark bump. Actual quantum calibration dropping from 10 minutes to under a second. Also: AI researchers might be automating themselves out of jobs at $4/hour, AWS shipped three ML releases in one day (because why not), and Meta built custom silicon proving you can still win with purpose-built chips. Read on.


🤖 AI & Agents

Anthropic’s Model Hardware Standard lets AI control lab equipment · Open spec enabling Claude and other agents to discover and operate physical devices through standardized APIs. Real results: Genentech automated protein assays with autonomous error recovery when robots failed mid-run. CMU cut serial dilution experiments from weeks to 8 hours of setup. QuEra improved quantum computer laser calibration from 58% to 99.3% success, dropping recovery time from 5-10 minutes to under a second. Partners: AWS (via Strands Robots library implementing MCP for device control), Automata, Hugging Face LeRobot, Universal Robots, Raspberry Pi. Research preview at modelhardwarestandard.com.
Read more · Can Your AI Agent Control Multiple Robots at Once?

Self-improving AI beats humans in 6 hours · Anthropic Fellow Chen Yueh-Han published research on Automated Alignment Researchers (AAR): AI systems that autonomously improve model alignment without human intervention. The system searches literature, proposes training methods, trains models in 30-minute intervals, and iterates. Result: beats experienced human researchers’ proposals within 6 hours, across all 10 alignment benchmarks tested. Cost: $4/hour. The humans cost $150/hour and are presumably rethinking their career choices. Early step toward recursive self-improvement where models improve their own training.
Read more

AWS AgentCore Evaluations launches · Framework-agnostic evaluation service for AI agents. Works with any agent emitting OpenTelemetry traces (Strands, LangGraph, OpenAI SDK, LlamaIndex, Claude SDK, Google ADK). Ships with built-in evaluators (GoalSuccessRate, Correctness, Helpfulness) plus custom LLM-as-a-judge support. Runs on-demand for CI/CD or continuously samples live traffic. Finally, standardized agent benchmarks.
Read more

Diagrid Catalyst 2.0 for AI agents · Adds durable and verifiable execution across ten frameworks (LangGraph, Microsoft Agent Framework, Google ADK, OpenAI Agents SDK, Claude Managed Agents, others). Durable execution: checkpoint-based recovery at the individual call level so interrupted runs resume without repeating completed work. Verifiable execution: cryptographic proof of what happened using SPIFFE-based identity. Solves the problem of costly agent failures when “a long multi-step run fails late in the sequence.”
Read more

AWS launches ARD spec for agent discovery · Agentic Resource Discovery (ARD) works like DNS for agents. Common protocol for describing agents, tools, and resources across clouds, on-prem, and SaaS. Local registries can federate without bilateral agreements. AWS Agent Registry implements it with MCP-native access and hybrid search. Apache 2.0 license.
Read more


🏆 Build with Agents — I’m Judging

Agents for Humans Hackathon: $40,000 in prizes, deadline Sep 14 · Build AI agents with Strands Agents SDK that handle repetitive tasks autonomously. Three tracks: Everyday Agents (daily life), Professional Agents (work productivity), Good Neighbor Agents (community impact). Grand prize: $10,000 cash + AWS feature. Each track: $5k/$3k/$2k for gold/silver/bronze. Requirements: open source (MIT/Apache), architecture diagram, demo video (max 5 min). Beginner friendly, 5,781 people registered already. I’m judging this one, so if you build something clever with agents, I’ll see it.
Register now


☁️ Cloud & Infrastructure

Meta built custom silicon for recommendation models · MTIA 300 is Meta’s first in-house training accelerator, optimized for ranking and recommendation (not LLMs). The problem: recommendation models spend significantly more time communicating between accelerators than computing, because “embedding tables can contain more than 99% of a recommendation model’s parameters.” Meta’s answer: two network chiplets with six custom 800 Gbps RDMA NICs each (1.2 TB/s total I/O), 16 dedicated message engines handling communication independently of compute, near-memory reduction hardware. Result: 3.9x reduction in total communication time vs equivalent GPU cluster, less than 0.5% compute degradation during concurrent operations (vs 20%+ on GPUs). Tested on 150-billion-parameter model across 40 accelerators.
Read more

DuckLabs joins AWS · The company behind DuckDB (the analytics database that’s absurdly fast) is joining Amazon. Projects stay open source under MIT license, governed by the independent DuckDB Foundation. This matters for AI workloads that need to crunch data inline without spinning up a warehouse.
Read more

AWS SageMaker Feature Store gets batch writes · New BatchWriteRecord API writes up to 25 records across feature groups in one call. ListRecords API enumerates record identifiers. If you’re managing ML features at scale, the batch write cuts latency.
Read more

Decathlon runs demand forecasting with Chronos-2 on AWS · Deployed Chronos-2 foundation model achieving 11-15 point forecast accuracy improvement. Weekly inference costs around $0.03 on CPU instances. Real production use case of foundation models for time series.
Read more

Salesforce hits Multi-AZ high availability with SageMaker Inference Components · Used SageMaker’s SchedulingConfig parameter to distribute models across Availability Zones for high availability compliance. If you’re deploying models under compliance requirements, this is the pattern.
Read more


✍️ From me this week

Deep in the agent memory series. This week: graph memory, observability, and DynamoDB vectors.

  • Graph Memory: When Vector Search Fails — Vector similarity finds memories but can’t connect them. Graph traversal does. Measured on multi-hop questions: 1/4 vs 4/4.
  • Observability for AI Agents with OpenTelemetry — How to debug agents in production when logs don’t cut it.
  • AI Agent Memory Part 2: DynamoDB Vector Search — Vectors inside your existing DynamoDB table. Single-digit ms, no separate database.

That’s it for this week. Claude can now control lab robots (what could go wrong?), AI researchers might be automating themselves out of jobs at $4/hour, and Meta built chips proving custom silicon still wins when you actually know your workload. If you’re building agents, the hackathon deadline is Sep 14. Come build something clever.

Dev.to · LinkedIn · GitHub · Twitter/X · Instagram · YouTube


¿Qué te pareció esta edición? Solo responde a este email.

📨 Tech & AI Weekly by Eli

The week in tech and AI: new models, tools, releases, and events worth knowing about. Curated so you stay informed without the noise. One email, every Friday.

No spam. Unsubscribe anytime. Double opt-in.

© 2026 Elizabeth Fuentes Leone. All rights reserved.

📨 Tech & AI Weekly by Eli

The week in tech and AI: new models, tools, releases, and what they mean for builders. Curated so you stay informed without the noise. One email, every Friday.

No spam. Unsubscribe anytime. Double opt-in.