Skip to Field Notes content

Agentic Engineering / Published from the workbench

Field Notes.

We study how capable models become accountable products: the harness, router, control plane, evidence, human decisions, and production consequences around them.

Weekly substantial note
Biweekly email

Complete English + Spanish editions

Archive

Start with the system.

7 issues · no filler

Issue 00714 min read

The Meta-Orchestrator Is the Workbench

What one anonymized, multi-slice regulated-industry push taught us about turning Orca, OMP, heterogeneous models, issue trackers, and CI into one accountable delivery system—and why adding agents is the easy part.

  • Agent orchestration
  • Orca ADE
  • OMP
  • Model routing
Read field noteLeer en español
Issue 00614 min read

Prism Arena Is Not a Leaderboard

Eleven models each received two recorded runs on the same ten animated 3D tasks. Four judges scored the selected artifacts. The useful result is not a universal winner. It is a record of what happened when models, output channels, schema validation, a shared renderer, frame capture, and a model panel were treated as one executable system.

  • Prism Arena
  • Systems benchmarks
  • Executable evaluation
  • LLM-as-a-judge
Read field noteLeer en español
Issue 00516 min read

The meter moved. It did not disappear.

On 20 July 2026, Claude Fable 5 became included in Max and Team Premium plans, capped at half the weekly usage pool. Four days later, Claude Opus 5 arrived on the API at half Fable 5's rate. What one week did to the economics of verification, and where a subscription lane must never carry a production dependency.

  • Fable 5
  • Opus 5
  • Model economics
  • Verification
Read field noteLeer en español
Issue 00417 min read

A model wave is a re-pricing event.

Five models and one access change landed between 16 and 24 July 2026. This is not a roundup. It is a routing note: which lanes got re-priced, which access boundary moved, what three of the four vendors did not publish in checkable form, and the policy we run because of it.

  • Model routing
  • Kimi K3
  • Gemini 3.6 Flash
  • Claude Opus 5
Read field noteLeer en español
Issue 00316 min read

Your evaluation sandbox is production infrastructure.

In July 2026 an internal OpenAI capability evaluation reached Hugging Face's production systems. Both companies have published accounts, and the accounts are not yet reconciled. The lesson does not depend on which one holds: the boundary around adversarial AI testing is a third-party risk surface, and it deserves production-grade controls. Written to the public record as of 25 July 2026.

  • Security
  • Agent evaluation
  • Containment
  • Incident response
Read field noteLeer en español
Issue 00216 min read

Autonomy is a rollback problem.

A field guide to expanding an agent's authority through bounded action spaces, exact approvals, failure drills, canary exposure, and recovery that works before production depends on it.

  • Agent systems
  • Autonomy
  • Guardrails
  • Evaluation
Read field noteLeer en español
Issue 00118 min read

The model is not the system.

A field guide to putting GPT-5.6 or Fable 5 inside an accountable operating system: one that routes deliberately, acts through a harness, keeps receipts, and ships work humans can verify.

  • Agent systems
  • GPT-5.6
  • Fable 5
  • Codex
Read field noteLeer en español

Weekly signal · biweekly email

Field-tested ideas. No content treadmill.

One substantial note when we have something worth showing: systems, receipts, prompts, and what failed. Confirm by email. Unsubscribe whenever you like.

Signup opens when the production Turnstile site key is configured.