Fetching from the wire…
Public story · 2026-08-04 · high
A lightweight proxy records real model calls from Codex and OpenClaw sessions, so agents train inside the harness itself instead of a simplified simulator.
Why now: Orchard lands as more coding agents get built directly on harnesses like Codex and OpenClaw, where a mismatched training environment shows up fastest.
Microsoft Research open-sourced Orchard, a framework that trains AI agents inside the real harnesses they'll actually run in, not a simplified stand-in.
Most agent reinforcement learning happens in a simplified environment. The mismatch shows up once the agent deploys into a real tool like Codex or OpenClaw. Orchard trains the agent end-to-end inside the harness itself, closing that gap before deployment.
The core piece is Orchard Env, a Kubernetes-based environment service that hands out reusable, isolated components for data collection, RL rollouts, and evaluation. It's built to work across domains without per-domain rework.
The mechanism is a lightweight proxy that sits inside the harness and records its own model calls as training data. Each rollout runs in its own container. That's what lets an agent train directly inside Codex or OpenClaw, instead of a mock version of either.
Microsoft is shipping three training recipes with the framework, along with training data and evaluation methods, per its research blog.
The bet here is that the gap between training environment and deployment environment, not a shortage of RL compute, has been capping agent quality. If that holds, agents trained on simplified stand-ins keep losing to harness-trained ones once both hit production, benchmarks aside.
Worth watching whether other labs open their harnesses to the same proxy trick, or whether this stays specific to Codex and OpenClaw.
Each link below shares sources, entities, or timing with this story.
Microsoft's July 23 release targets a genuine gap: harness-based agents like Claude Code and Codex drive multi-turn reasoning, tool use, and external system access but were hard to train end-to-end with standard open RL infrastructure. The trick is decoupling training from inf...
A paper from Xiao Yu, Baolin Peng, and Ruize Xu makes a claim that seems obvious once stated and is genuinely new as a training methodology: modern agents are inseparable from their inference harnesses, so training them in stripped-down RL sandboxes produces a train/serve mism...
July 22's board is plumbing, not apps: Kastra (199 votes) sells runtime authorization for Claude, Cursor, Codex, and OpenClaw with policy enforcement against prompt injection and unauthorized tool calls, while box (183 votes) sells Ubuntu VMs with SSH for agents at $0.036/hr (...
A category has formed around one job, watching and steering many concurrent coding agents from a single pane. AionUi (TypeScript) markets a 24/7 "Cowork" app spanning OpenClaw, Hermes, Claude Code, Codex, OpenCode and 20+ more CLI agents; agent-of-empires (Rust) offers TUI and...
QM went up under MIT license. Created July 29. As of the GitHub API check: 8,420 stars, 887 forks. Five days. YC uses it internally across accounting, legal, events, and engineering, including to build QM itself. Every employee and every Slack room gets its own scoped memory,...
I check Product Hunt maybe once a week and usually regret it. Today's board is worth reading as market structure. The July 30 leaderboard: SKI at 277 upvotes (free voice input for Claude Code and Codex). AI Search Console at 249 (prompt analytics and citation mapping). Memmy A...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.