The meter moved. It did not disappear.
On 20 July 2026, Claude Fable 5 became included in Max and Team Premium plans, capped at half the weekly usage pool. Four days later, Claude Opus 5 arrived on the API at half Fable 5's rate. What one week did to the economics of verification, and where a subscription lane must never carry a production dependency.
01 / The meter
Verification is the work that lives at the margin.
Generation runs once and produces the thing you wanted. Verification is optional, repeatable, and invisible when it succeeds. Under per-token billing, every additional check is a marginal cost with no visible payoff, and marginal costs with no visible payoff are where budget pressure settles. That is an argument about incentives, not a survey of the industry; we have no interview data and will not pretend to. What we can report is our own ledger. When we audited our pipeline this quarter, the checks we were not running clustered exactly where the incentive analysis predicts: the second opinion, the adversarial re-read, the eval suite that ran before releases instead of on every commit. Nobody had decided that in a meeting. It had settled there.
The absolute numbers frame the decision. Claude Fable 5 is priced at 10 dollars per million input tokens and 50 per million output on the API. Claude Opus 5, announced 24 July 2026, is 5 and 25. The overhead is smaller and easier to forget: on Opus 5, a tool-equipped request carries 286 tokens of tool-use system prompt at tool_choice auto or none, 406 at any or tool, and another 325 input tokens if the bash tool is attached. Claude Managed Agents bills 0.08 dollars per session-hour on top of Opus 5 token rates. The bill was never only tokens, and a fan-out of small verification calls pays the fixed overhead on each one.
Four practices sit closest to that margin. Redundant verification: the same check through two independent paths. Second opinions: a different model, with a different training cutoff and different failure modes, reviewing the same artefact. Adversarial self-review: the system instructed to attack its own output rather than summarise it. Continuous evaluation: the suite on every commit rather than before a launch. Each is cheap in isolation. Each is expensive as a default, on every change, indefinitely, which is why each is the natural first thing to quietly not do.
- Generation is billed once; verification is billed every time you take it seriously.
- Tool-equipped calls carry fixed per-call overhead before any work happens; fan-out pays it repeatedly.
- Audit your own pipeline before believing the incentive argument. Ours confirmed it; yours is the one that matters.
- Skipped verification never appears as a line item, only as an absent habit.
02 / 20 July
Read the terms before you read the headline.
Effective 20 July 2026, Anthropic included Claude Fable 5 at no extra cost on Max plans, premium Team seats, and premium seat-based Enterprise seats. The help centre states the allocation plainly: up to 50 percent of weekly usage limits may go to Fable 5 at no extra cost. It is not universal. Pro and standard Team seats are excluded; on those plans Fable 5 is not included in the plan's usage limits and is paid for with usage credits. The inclusion names premium seat-based Enterprise seats specifically, not standard ones. On the transition credit, the sources disagree: the official @claudeai account, quoted by Simon Willison on 18 July 2026, said excluded users would receive a one-time 100 dollar credit, and PCWorld reported the same figure; Anthropic's help centre says only that there is a one-time credit and states no amount.
The change is dated, not promised. The help centre article gives a start date, 20 July 2026, and no end date, and it does not use the word permanent. Several outlets read the absence of an end date as permanence; that is their inference, not the company's text. A team planning capacity on a commitment when it holds only a start date is making a forecasting error. For scale: Max starts at 100 dollars per month, Team Premium is 100 dollars per seat annual or 125 monthly, Pro is 17 annual or 20 monthly, Team Standard is 20 or 25. The inclusion tracks the expensive seats. It also did not replace nothing: a prior free-access promotion ran until 19 July 2026 at 11:59:59 PM PT, after which Fable 5 remained available under the new terms.
Now the mechanics. The 50 percent figure is a ceiling on a share, not an addition. The help centre is explicit that use of other models draws from the same usage limits and that you can never exceed the weekly limit. The pool is shared: everything else that seat does spends from the same allowance. The change gives you permission to put half the pool through the stronger lane; it does not give you a bigger pool. The article describes the 20 July change and says nothing about Claude Opus 5, and we found no published statement on how the model that became the Max default four days later interacts with the allowance. On the general wording it draws from the same pool, but that is our reading, not a documented one.
Two more gaps. First, no plan tier has published weekly limit values in any unit we could find: not messages, not tokens, not hours. You cannot budget against a number nobody published; you can only measure your own exhaustion. The-decoder reported that the 50 percent applies to limits already reduced by roughly a third when the bonus phase ended; we could not verify that against a primary source and flag it as their reporting. Second, there is no low-intensity Fable 5 request to fall back on. Adaptive thinking is always on; the disabled thinking type is unsupported and raw chain-of-thought is never returned. And if your token intuitions were formed on older models, recalibrate: Fable 5 uses the tokenizer introduced with Claude Opus 4.7, and Anthropic's documentation notes that compared to models before Opus 4.7, the same text produces roughly 30 percent more tokens.
- Included: Max, premium Team seats, premium seat-based Enterprise seats. Excluded: Pro and standard Team, which pay with usage credits.
- The allocation is up to 50 percent of the weekly limit, shared with every other model on the seat. Not additive.
- The help centre article states a start date and no end date; it does not use the word permanent.
- Reported as a 100 dollar one-time credit by @claudeai and PCWorld; the help centre states no figure.
- No plan's weekly limit is published as a number, in any unit. Measure it or guess.
- Thinking cannot be switched off on Fable 5; there is no cheap request shape.
Fast lane
Extraction · formatting · deterministic checks
Balanced lane
Default implementation · routine research
Frontier lane
Architecture · ambiguity · security · synthesis
Human gate
Consequential action · approval · acceptance
03 / 24 July
Four days later the metered lane got cheaper.
On 24 July 2026 Anthropic announced Claude Opus 5 at 5 dollars per million input tokens and 25 per million output. That is half Fable 5's rate and the same price as its predecessor Opus 4.8, so it is a capability change, not a price cut. Anthropic positions it as coming close to Fable 5's frontier intelligence at half the price and recommends it as the default starting model; it became the default on Claude Max and the strongest model available on Claude Pro. The hierarchy is unchanged at the top: Fable 5 remains, per Anthropic's own documentation, their most capable widely released model, and Opus 5 remains behind Claude Mythos 5 on cybersecurity and biology work.
On capability, be careful what you repeat. Both launches present their benchmark results as charts rather than text tables, and we could not extract a single absolute figure for either model from any source we checked. What exists is Anthropic's own comparative wording: that Opus 5 surpasses all other models on Frontier-Bench v0.1, comes within 0.5 percent of Fable 5's peak CursorBench 3.2 score at half the cost, and surpasses Fable 5's best OSWorld 2.0 result at just over a third of the cost. Those are the vendor's claims about relative position. A comparison you cannot reproduce is a claim, not a measurement, and we cite it as one.
The metered lane also carries controls the plan surface does not document. Effort is set through output_config with five levels, low through max, defaulting to high on the Claude API and in Claude Code; thinking cannot be disabled at xhigh or max, and Anthropic's docs note that on Opus 5, changing effort does not reliably shorten the visible response, so prompt for length instead. Prompt caching prices the redundancy directly: cache hits and refreshes at 0.50 dollars per million input tokens against a 5 dollar base, five-minute cache writes at 6.25, one-hour writes at 10. The Batch API halves the rate again to 2.50 and 12.50. Long context bills at standard rates with no surcharge above 200k tokens on Claude 4.6 and later, and automatic fallbacks shipped as an API beta alongside the launch. We found nothing documenting an equivalent set of dials on subscription seats.
Hold the week in one frame. On 20 July, strong-model verification became a quota draw on expensive seats. On 24 July, a metered model arrived at half the frontier rate with caching, batch, and effort controls. Any verification practice hard-wired to one lane's price sheet needed re-deciding within the week. A practice built on a rubric, a deterministic gate, and a swappable lane needed a config change.
- Opus 5: 5 and 25 dollars per million tokens. Same price as Opus 4.8, half of Fable 5.
- Benchmark results for both launches are charts; we could not extract one absolute figure. Cite the comparative claims as claims.
- Effort, caching, batch, and automatic fallbacks are API-side levers; no equivalent is documented on plan seats.
- Route by rule, not by price sheet. The price sheet changed twice in five days.
04 / The panel
What we run, with the edges showing.
Before a pull request opens in our pipeline, a change passes a deterministic verifier and a multi-model review panel. The verifier is the gate: build, type check, tests, lint, migration checks, and the project's own assertions. It is the only thing permitted to fail a change. The panel is advisory. Several model lanes review the same diff against the same rubric with the same repository context, and they write findings. They do not merge, they do not approve, and they cannot mark their own acceptance criteria as verified.
The signal we route on is disagreement. Where the panel converges, a human skims. Where it splits, a human reads carefully. That is a working rule, not a validated finding: we have not measured whether splits predict defects better than any single reviewer's severity ranking, and until we publish those run records, treat the heuristic as our bet rather than our receipt. The honest version of the economics is similar. A panel is only worth running if redundancy is cheap, and cheap has two paths: a subscription lane whose capacity is unpublished, or a metered lane priced for exactly this shape of work. Within one model lane, a panel is unusually cache-friendly, because every reviewer reads the same context and differs only in rubric. Fable 5 retains the 90 percent input token discount for prompt caching; Opus 5's cache-hit rate is a tenth of its base input price. Whether a cache entry can be shared across different models is not documented, so budget each lane's context write separately, and remember that every reviewer pays its own output tokens regardless.
One counterargument deserves the floor: tokens may not be your binding constraint. Every reviewer you add produces findings a human must triage, and false positives spend the scarcest resource in a small team, which is attention. Our mitigations are caps and buckets: a ceiling on reviewers per run, findings required to carry a file, a line, and a concrete failure scenario, and anything without a reproduction path diverted to a low-signal bucket a human can ignore. If the panel makes review slower without catching a defect class your gate misses, the panel is theatre. Measure that, not the token bill.
Two operational cautions from running this. Staleness: Fable 5's reliable knowledge cutoff is January 2026, Opus 5's is May 2026, and a reviewer months behind on a library's API will confidently flag correct code as wrong. Choose reviewing lanes on what they know, not only on what they cost. Refusals: Fable 5 ships safety classifiers that can decline a request, and a refusal arrives as an HTTP 200 with stop_reason set to refusal; requests refused before output are not billed. A harness that treats HTTP 200 as success will record a refusal as a passing review. Assert on stop_reason, and count a refusal as a non-review that visibly reduces coverage.
- The deterministic verifier is the gate; the model panel is advisory. Never the reverse.
- Disagreement routing is our working heuristic, not a validated finding. We say so until we publish run records.
- Cache the shared context within each lane; cross-model cache reuse is not documented, so do not assume it.
- Cap reviewers and quarantine findings without a reproduction path. Human attention is the binding constraint.
- Assert on stop_reason, not HTTP status. A refusal is a 200 and must reduce reported coverage.
05 / What a subscription is not
You swapped a cost risk for an availability risk.
A seat subscription is not an API service level agreement, and the record of this particular model makes the difference concrete. Between 9 June and 20 July 2026, Anthropic revised Fable 5's plan access at least four times: a free window from 9 to 22 June, a switch to usage credits on 7 July, an extension on 12 July, and the inclusion change on 20 July. Anthropic said why: demand for Fable had been challenging to predict, which is why access was extended several times as additional capacity was secured. Read that as an engineering fact. The terms of the lane are downstream of capacity, and capacity is not a commitment you hold.
Availability has a harder precedent than terms drift. Fable 5 was forced offline on 12 June 2026 by a United States export-control directive. Controls were lifted and the model returned globally on 1 July 2026. Nineteen days, during which no retry policy, no fallback configuration, and no amount of budget on your side would have restored that specific model. It was not a rate limit or an outage. It was regulation, and a dependency map that has never considered a regulatory removal has a blind spot the size of June 2026.
One constraint travels with the model rather than the plan. Fable 5 is designated a Covered Model with 30-day data retention and is not available under a zero data retention agreement. If a client contract requires ZDR, Fable 5 is off the table for that engagement on every lane, subscription or API. That has nothing to do with pricing and will not surface in any cost review, which is exactly why it belongs in the routing rule rather than in someone's memory.
- At least four term revisions between 9 June and 20 July 2026.
- Anthropic's stated cause: demand was hard to predict; access extended as capacity was secured.
- Nineteen days offline in June 2026 under a US export-control directive. Regulation is not retryable.
- Fable 5 carries mandatory 30-day retention and cannot run under ZDR, on any lane.
06 / Where the boundary sits
Advisory, retryable, asynchronous, human-gated.
The rule we use has four conditions, and work must satisfy all four to live on a subscription lane. Advisory: nothing merges or ships on its say-so. Retryable: a failed run can simply run again later. Asynchronous: nobody is holding a latency budget against it. Human-gated: a person accepts or rejects the output before it has any effect. The conditions attach to how work is run, not to what it is called. A pull request review panel run as ours is, pre-merge but advisory and human-gated, qualifies; the same panel wired to block merges does not, and a review a customer is waiting on does not either. Eval generation, spec critique, red-teaming, migration planning, and documentation drafts usually qualify as we run them, but check the conditions, not the category.
The other side of the line is anything a client's user waits on, anything covered by a support commitment, anything with a customer-visible failure mode, and anything governed by a data-processing term the subscription does not carry. Those stay on the API, where the controls are documented and contractual: pinned model identifiers, claude-opus-5 being a dateless string but a pinned snapshot rather than a moving alias; automatic fallbacks, shipped as an API beta; batch pricing for work that tolerates latency; and a cost that lands on an invoice where a finance function can see and challenge it.
The failure mode to watch is not a bad decision at the boundary. It is an undeclared dependency drifting across it. A panel that starts as a helpful habit becomes expected within a quarter, and then a seat subscription with unpublished limits and terms that moved four times in six weeks is sitting on your delivery path with nobody having written that down. The fix is unglamorous: declare the lane in the same place you declare your CI provider, give it a documented spillover to a metered lane with a pinned model, and force the spillover to fire on a schedule so you know it works.
Which lane should host the panel by default? We are not going to pretend this article settles it. The subscription lane's marginal price is zero but its capacity is unpublished and its terms are five days old. The metered lane has published prices, caching, batch, and effort controls, and Opus 5 moved its economics substantially on 24 July. The comparison depends on your diff sizes, your cache hit rates, and your run volume, none of which we can know for you. What we can say is that the four-condition rule and the spillover survive either answer, and a practice that needs re-litigating every time a price sheet moves was never a practice.
- Subscription lane: advisory, retryable, asynchronous, human-gated. All four, judged by how the work runs.
- Metered lane: latency budgets, uptime commitments, customer-visible paths, contractual data terms.
- Declare the subscription lane as a dependency the day the panel becomes expected.
- Exercise the spillover on a schedule; an untested fallback is a comment.
- The lane choice is measurable per team. The routing rule is the part that transfers.
07 / What to do now
Build the gate, then spend the cheap lane on redundancy.
Start with the ledger. Write down every verification step your pipeline does not run, and next to each one the reason. Some reasons will be engineering judgement. Some will be a cost estimate from a different price sheet that nobody has revisited. Recompute every too-expensive entry against the rates in this article before accepting it, because those rates changed twice in the week before publication. That list, not a benchmark chart, tells you what a cheap verification lane is worth to your team.
Then build in the right order. The deterministic verifier comes first; it is what makes an advisory panel safe to run at all, because a panel that can block a merge without a reproducible criterion makes delivery slower and no more correct. Once the gate is the only thing that can fail a change, add reviewers, cap them, and treat their disagreement as a routing signal for human attention while you collect the run records that will tell you whether the signal is real.
Treat the quota as a budget you have to discover. The weekly limits are not published, so learn them from your own side: log what you route through the lane, log when it exhausts, project exhaustion from observed burn, and decide in advance what degrades when the allowance runs out. The degradation should be a declared spillover to a metered lane with a pinned model, visible in the run record, never a silent drop to a single reviewer that nobody notices until a defect ships.
And date everything. The terms described here took effect on 20 July 2026 and were five days old when this was written; they had already changed at least four times in six weeks, and the model landscape moved again on 24 July. Re-verify against the primary sources below before any plan depends on a claim in this article. We hold our own writing to the same rule we hold a model's findings to: a claim without a checkable source is a liability, whoever produced it.
- List the verification you skip; recompute every cost-based reason at current rates.
- Ship the deterministic gate before the model panel, never the reverse.
- Keep models advisory: no merge authority, no self-certification of criteria.
- Learn the quota from observed exhaustion; alert on projected burn, not after the fact.
- Declare the spillover lane with a pinned model ID and fire it on a schedule.
- Re-verify plan terms before any capacity plan depends on them. They are days old.
08 / Production missions
Prompts that specify a whole mission.
These are not idea starters. Each is an operating contract for GPT-5.6 or Fable 5: outcome, boundaries, architecture, model policy, phases, acceptance tests, evidence, release, and stop conditions.
Mission 01 · Verification
Deterministic gate with an advisory model panel
A pre-PR pipeline where a reproducible verifier is the only thing that can fail a change, and several model lanes add redundancy without acquiring authority.
Deterministic gate first · subscription and metered model lanes as panel members
Mission 01 · Verification
Deterministic gate with an advisory model panel
A pre-PR pipeline where a reproducible verifier is the only thing that can fail a change, and several model lanes add redundancy without acquiring authority.
Deterministic gate first · subscription and metered model lanes as panel members
Full operating contract
MISSION: BUILD A DETERMINISTIC GATE AND AN ADVISORY MULTI-MODEL REVIEW PANEL Act as the release engineer and verification owner for a small team. Build the checks that run on a change before a pull request opens. The deterministic verifier is the gate. The model panel is advisory and must never acquire merge authority, either by design or by drift. INPUTS I WILL PROVIDE - Repository, build and test commands, type checker, linter, migration tooling, and current CI configuration. - Available model lanes with their identifiers, cost basis, and whether each is metered or drawn from a subscription quota. - Review rubric, historical defect classes, and any changes that previously shipped broken. - Data handling constraints, including any zero data retention requirement on client work. OPERATING CONTRACT 1. The deterministic verifier runs first and short-circuits. The panel never reviews a change that does not build. 2. Only the deterministic verifier may fail a change. Panel findings are advisory and are attached, never enforced. 3. No model may mark its own acceptance criterion as verified, approve a change, or open a pull request. 4. Pin every model identifier. Record the identifier in the run record alongside the finding. 5. Assert on the response stop reason, not on HTTP status. A refusal can arrive as a success status with a refusal stop reason and must be recorded as a non-review, not a pass. 6. Store observable traces and concise findings. Never store or request hidden reasoning. 7. Route any lane that cannot satisfy the engagement's data retention terms out of the panel entirely, with the exclusion recorded. GATE DESIGN - Build, type check, unit and integration tests, linter, formatter, migration up and down, and project-specific assertions. - Every gate check is reproducible from a clean checkout with no model in the loop. - Failures report the exact command, exit code, and first actionable output line. - The gate emits a machine-readable result that the panel receives as context. PANEL DESIGN - Assemble the shared context once: diff, surrounding files, gate result, rubric, and repository conventions. - Within each model lane, write the shared context to that lane's prompt cache once and have its reviewers read from it. Do not assume a cache entry is shared across different models; budget each lane's context write separately. - Give each reviewer a distinct assignment: correctness, security and authorization, contract and interface stability, test adequacy, operational and rollback impact. - Require every finding to carry a file, a line range, a severity, a reproduction or a concrete failure scenario, and a confidence. - Divert findings without a reproduction path into a separate low-signal bucket rather than the main list. - Enforce a ceiling on reviewers per run. Human triage capacity, not token price, sets the ceiling. DISAGREEMENT ROUTING Compute agreement across reviewers per finding location. Where reviewers converge on no findings, mark the change for a skim. Where they converge on a finding, surface it first. Where they split on the same location, escalate for close human reading and record the split as the reason. Treat this routing as a hypothesis: log outcomes so that, after enough runs, you can test whether splits actually predict defects. Do not resolve disagreement with a majority vote or a judge model in the first version. QUOTA AND COST DISCIPLINE - Record per run: lane, pinned model identifier, whether the lane was metered or quota-drawn, wall time, and token counts where the lane exposes them; where it does not, record request counts and payload sizes instead. - Enforce a per-run ceiling on total reviewers and a per-day ceiling on panel invocations. - When a quota-drawn lane is unavailable or exhausted, spill to a declared metered lane with a pinned identifier and record the substitution in the run record. ACCEPTANCE TESTS - A change that fails the gate never reaches the panel and consumes no model budget. - A panel finding cannot block a merge through any code path, including future ones. Prove it with a test, not a convention. - A refused model response is recorded as a non-review and visibly reduces panel coverage for that run. - Removing one reviewer changes coverage reporting rather than silently passing. - Two runs over an identical diff with pinned lanes produce identical gate results and comparable panel coverage. - Cache usage is observable in the run record for every metered lane that reports it. - A lane excluded for data retention reasons cannot be selected by fallback. DELIVERABLES Gate configuration, panel orchestrator, per-lane cache strategy, rubric set, run record schema, disagreement routing logic, spillover policy, and a demonstration on three historical changes: one that should fail the gate, one clean, and one where reviewers legitimately disagree. End with the findings the panel missed on those three and what the gate would need to catch them instead.
Mission 02 · Model operations
Quota-aware routing with a declared spillover
A routing policy that treats a subscription allowance as a depletable budget of unknown size with its own client-side telemetry, and that degrades to a metered lane on purpose rather than by accident.
Subscription lane as primary for advisory work · pinned metered API lane as spillover
Mission 02 · Model operations
Quota-aware routing with a declared spillover
A routing policy that treats a subscription allowance as a depletable budget of unknown size with its own client-side telemetry, and that degrades to a metered lane on purpose rather than by accident.
Subscription lane as primary for advisory work · pinned metered API lane as spillover
Full operating contract
MISSION: MAKE A SUBSCRIPTION QUOTA A FIRST-CLASS BUDGET IN THE ROUTER You are building the routing policy for a team that has started running advisory work on a subscription lane whose weekly limits are not published. Treat the allowance as a depletable resource with unknown capacity, shared across every model on the seat, and reset on a schedule you do not control. All measurement is client-side: assume the vendor exposes no draw telemetry. The router must make exhaustion boring. INPUT CONTRACT - Current lanes with identifiers, billing basis, allocation caps where documented, and known unavailability history. - Work classes with risk level, latency tolerance, retryability, and whether a human gates the output. - Any contractual data terms that exclude specific models from specific engagements. - Existing telemetry, cost ledger, and alerting destinations. CLASSIFICATION RULE A work class may default to a subscription lane only if it is advisory, retryable, asynchronous, and human-gated. All four must hold, judged by how the work actually runs, not by its category name. Encode this as a check on the work class definition, not as guidance in a document. Any class failing one condition defaults to a metered lane with a pinned identifier, and the router refuses to route it to a subscription lane even under manual override without a recorded human decision. QUOTA MODEL - Model the allowance as an estimated capacity with a confidence interval, learned from observed exhaustion events rather than from a published number. - Track what you send: requests, payload sizes, and token counts where the lane reports them, per work class, per day, per seat. Attribute draw from every model sharing the pool, not only the frontier lane. - Account for the tokenizer in use when converting text volume to expected draw. Do not carry forward token estimates calibrated on an older tokenizer. - Surface projected exhaustion date at current burn, and alert on the projection crossing the reset boundary rather than on absolute usage. SPILLOVER POLICY For every work class define, in machine-readable form: primary lane, spillover lane with pinned identifier, the trigger conditions for spillover, the maximum spend the spillover may incur before it in turn degrades, and what the system does when both are unavailable. Spillover must be visible in the run record. A silent substitution is a defect. AVAILABILITY HANDLING - Distinguish rate limiting, transient error, refusal, model-level unavailability, and regional or regulatory unavailability. Each has a different correct response. - Treat prolonged model-level unavailability as a first-class scenario with an explicit runbook, not as an extended retry. - Where the provider offers automatic fallback features, decide explicitly whether to use them or to keep fallback in your own control plane, and record the reason. - Never let a retry loop convert an availability failure into quota exhaustion. IMPLEMENTATION ORDER A. Define work classes and the four-condition eligibility check with tests. B. Add client-side draw telemetry and attribution before changing any routing decision. C. Run in shadow mode: log the decision the policy would make, keep current behaviour. D. Enable the policy for one advisory class and observe a full quota reset cycle. E. Add spillover, then rehearse it by forcing the primary lane unavailable in a working week. F. Only then widen to further classes. ACCEPTANCE TESTS - A class that fails any of the four conditions cannot be routed to a subscription lane, including through fallback and including through manual override without a recorded approval. - Quota exhaustion produces a declared spillover and an alert, never a silent capability reduction. - A forced primary-lane outage is absorbed within the declared spillover budget and appears in the run record. - Retry storms cannot consume more than a bounded share of the estimated remaining allowance. - The projected exhaustion estimate improves measurably after the first observed exhaustion event. - Every routing decision is reconstructable from the run record without access to the router source. DELIVERABLES Work class definitions, eligibility tests, quota estimator, attribution telemetry, spillover policy as machine-readable configuration, availability runbook, shadow-mode comparison report, and a rehearsed outage drill with its receipt. State plainly which capacity numbers you had to estimate because no vendor documentation published them.
Mission 03 · Engineering practice
Audit what per-token pricing quietly removed
A written ledger of every verification step the team stopped doing for cost reasons, priced under current rates, with a ranked restart plan.
Cheap lane for inventory and extraction · deliberate lane for the cost model and ranking
Mission 03 · Engineering practice
Audit what per-token pricing quietly removed
A written ledger of every verification step the team stopped doing for cost reasons, priced under current rates, with a ranked restart plan.
Cheap lane for inventory and extraction · deliberate lane for the cost model and ranking
Full operating contract
MISSION: FIND THE VERIFICATION YOUR PRICING MODEL DELETED Act as an engineering practice auditor. Reconstruct which checks this team does not run, separate genuine engineering judgement from unexamined cost avoidance, and price the difference under current model rates. Produce a ranked restart plan, not a lament. INPUTS I WILL PROVIDE - CI configuration, pre-commit hooks, review checklists, eval suites and their trigger conditions, and any runbooks. - Incident and defect history, including changes that shipped broken and the check that would have caught each one. - Current model lanes, their rates, caching behaviour, batch availability, and any subscription allowance. - Team size, review capacity, and the real constraint on how much a human can read per week. INVENTORY PHASE For every verification step that exists, record: what it checks, when it runs, what triggers it, what it costs per run in tokens and wall time, and whether it is deterministic or model-mediated. Then, for every step that does not exist but plausibly should, record the same fields as an estimate and mark it clearly as an estimate. ATTRIBUTION PHASE For each absent step, assign exactly one primary reason from: not valuable, superseded by another check, not technically feasible here, too slow for the workflow, too noisy, or too expensive. Cite evidence for the reason from configuration history, commit messages, or a named person's recollection marked as such. Do not accept too expensive as a reason without an actual figure. Recompute that figure under current rates before accepting it. COST MODEL - Price each candidate step per run and per week, separating cached from uncached input, and state every assumption (diff size, context size, cache hit rate) explicitly as an assumption. - Price it three ways: naive metered, metered with caching and batching applied, and drawn from a subscription allowance. - For the subscription case, express cost as a share of the weekly allowance rather than in currency, and state explicitly that the allowance size is estimated if it is not published. - Include fixed per-call overhead for tool-equipped requests, and any per-session runtime charges. Do not model tokens alone. - Weigh the human side: findings produced per run and the triage minutes they consume. A step that is cheap in tokens and expensive in attention is not cheap. RANKING Rank candidate steps by defect classes historically caught per unit of cost, then by review time saved, then by how gracefully the step degrades when its lane is unavailable. Steps that are advisory, retryable, asynchronous, and human-gated rank higher for subscription-lane placement. Steps on a customer-visible path rank higher for metered placement regardless of cost. OUTPUT - A ledger table: step, status, reason for absence, evidence, cost under each of the three pricing models, historical defect classes addressed, and recommended lane. - A shortlist of at most five steps to restart this month, each with its trigger, its lane, its budget, and the specific defect class it targets. - A list of steps you recommend leaving off, with the reason, so the decision is recorded rather than re-litigated every quarter. ACCEPTANCE TESTS - Every too-expensive attribution carries a recomputed current figure or is reclassified. - Every recommended step names a defect class from real history, not a hypothetical one. - Cost figures distinguish cached from uncached input and state the assumed cache hit rate as an assumption. - Subscription-lane recommendations state the allowance share and flag that the allowance is estimated. - No recommended step places a customer-visible dependency on a subscription lane. - The ledger is reproducible: a second auditor with the same inputs reaches the same shortlist. DELIVERABLES The ledger, the shortlist with budgets and triggers, the declined list with reasons, and a one-page note on what changed in your cost assumptions since they were last examined. Date the whole document and name the rates it was computed against, because those rates move.
09 / En español
Resumen en español
El 20 de julio de 2026 Anthropic incluyó Claude Fable 5, sin costo adicional, en los planes Max, en los asientos Team Premium y en los asientos premium de Enterprise, con un tope: hasta 50% del límite semanal de uso. No es un límite más grande; es una porción del mismo bote, compartido con todos los demás modelos del asiento. Pro y Team Standard quedaron fuera y ahí Fable 5 se paga con créditos de uso; la inclusión menciona solo los asientos Enterprise premium, no los estándar. Sobre el crédito de transición las fuentes difieren: la cuenta oficial @claudeai habló de 100 dólares por única vez; el centro de ayuda de Anthropic confirma el crédito pero no publica el monto.
El artículo de ayuda fija una fecha de inicio y no fija una de término, y no usa la palabra permanente; esa lectura es de la prensa, no del texto de la empresa. Entre el 9 de junio y el 20 de julio los términos cambiaron al menos cuatro veces, y el 12 de junio una directiva de control de exportaciones de Estados Unidos sacó el modelo de línea hasta el 1 de julio: diecinueve días en los que ninguna decisión de ingeniería del lado del cliente habría servido. Además, ningún plan publica el valor de su límite semanal en ninguna unidad. Todo lo de esta nota está fechado al 25 de julio de 2026 y hay que volver a verificarlo antes de depender de ello.
Cuatro días después, el 24 de julio, salió Claude Opus 5 en la API a 5 y 25 dólares por millón de tokens: la mitad de la tarifa de Fable 5 y el mismo precio que su antecesor Opus 4.8, o sea un salto de capacidad, no una rebaja. Ninguno de los dos lanzamientos publicó cifras absolutas de benchmarks que pudiéramos extraer; lo que existe son las afirmaciones comparativas del propio proveedor, y las citamos como afirmaciones. En una semana la economía de la verificación cambió dos veces, y esa es la razón para no amarrar la práctica a una tarifa.
Nuestro uso es concreto: antes de abrir un pull request, un verificador determinista y un panel de revisión con varios modelos. El verificador es lo único que puede reprobar el cambio; el panel opina y ya. Ruteamos la atención humana por el desacuerdo entre revisores, y lo decimos con honestidad: es una regla de trabajo, todavía no una conclusión medida. La frontera que usamos tiene cuatro condiciones para el carril de suscripción: asesor, reintentable, asíncrono y con una persona de por medio, las cuatro juntas. Lo que tiene presupuesto de latencia, compromiso de disponibilidad, cara visible al cliente u obligación contractual de datos se queda en la API, con modelos fijados y una ruta alterna que se ensaya. Ojo aparte: Fable 5 retiene datos 30 días y no opera bajo retención cero; si el contrato exige ZDR, queda fuera en cualquier carril.
- Haz la lista de las verificaciones que no corres y recalcula cada motivo de costo con las tarifas vigentes.
- Primero el verificador determinista, después el panel de modelos. Nunca al revés.
- Ningún modelo aprueba su propio criterio ni hace merge.
- El límite semanal no está publicado: apréndelo de tu propio consumo y alerta por proyección de gasto.
- Declara el carril alterno medido con modelo fijado y ensáyalo con calendario.
- Los términos tienen días de vigencia y cambiaron cuatro veces en seis semanas: verifica antes de planear con ellos.
10 / Source trail
Primary, fresh, inspectable.
Product details move quickly. These sources were checked for this issue on July 25, 2026. Re-verify before changing production policy.
- 01Claude Fable 5 on your planAnthropic SupportPrimary source for the 20 July 2026 change: inclusion on Max, premium Team, and premium seat-based Enterprise seats; the 50 percent allocation; exclusion of Pro and standard Team; the shared weekly pool; and the one-time credit with no stated amount.
- 02Introducing Claude Fable 5 and Claude Mythos 5AnthropicPrimary source for Fable 5 pricing, 9 June 2026 availability, always-on adaptive thinking, refusal behaviour, 30-day retention and Covered Model status, and supported API features including effort and task budgets.
- 03Models overviewAnthropicPrimary source for model identifiers, context windows, knowledge cutoffs, the Opus 4.7 tokenizer note, and the description of Fable 5 as Anthropic's most capable widely released model.
- 04Redeploying Fable 5AnthropicPrimary account of the June 2026 export-control suspension and the 1 July 2026 return to global availability.
- 05Claude FableAnthropicSource for the 90 percent prompt caching input discount and the US-only inference multiplier.
- 06Introducing Claude Opus 5AnthropicPrimary source for the 24 July 2026 launch, the half-the-price positioning, default status on Max and strongest-on-Pro, and the comparative benchmark claims, presented as charts from which we could extract no absolute figures.
- 07PricingAnthropicPrimary source for Opus 5 token rates, prompt cache write and hit rates, Batch API rates, tool-use overhead, Managed Agents session billing, and the absence of a long-context surcharge.
- 08EffortAnthropicPrimary source for the five effort levels on Opus 5, the default of high on the Claude API and Claude Code, and the constraint on disabling thinking at xhigh and max.
- 09Claude plan pricingAnthropicSource for Pro, Max, Team Standard, Team Premium, and Enterprise seat prices as of July 2026.
- 10Claude will make Fable 5 permanentSimon WillisonSecondary source quoting the official @claudeai statement of 18 July 2026, including the 100 dollar one-time credit and Anthropic's explanation of the staged rollout. The original X post could not be fetched directly.
- 11Fable will stay in Claude plans, but not for everyonePCWorldSecondary corroboration of the 100 dollar one-time credit figure that the help centre does not state.
- 12Anthropic slashes Claude Fable 5 limits in Max and Team PremiumThe DecoderIndependent reporting that the 50 percent allocation applies to already-reduced limits. Cited as reporting; we could not verify the reduction against a primary source.
- 13Anthropic launches Opus 5TechCrunchIndependent confirmation of the 24 July 2026 Opus 5 launch date.
Weekly signal · biweekly email
Field-tested ideas. No content treadmill.
One substantial note when we have something worth showing: systems, receipts, prompts, and what failed. Confirm by email. Unsubscribe whenever you like.
