Skip to Field Notes content
Field Note 00417 min read

A model wave is a re-pricing event.

Five models and one access change landed between 16 and 24 July 2026. This is not a roundup. It is a routing note: which lanes got re-priced, which access boundary moved, what three of the four vendors did not publish in checkable form, and the policy we run because of it.

01 / The category error

A release wave is a re-pricing event, not a shopping list.

Between 16 and 24 July 2026, five models shipped and one access boundary moved. Moonshot AI released Kimi K3 in mid-month, Alibaba previewed Qwen3.8-Max on 19 July and announced Qwen-Image-3.0 around the same days, Google shipped Gemini 3.6 Flash on 21 July, and Anthropic launched Claude Opus 5 on 24 July. The sixth event is different in kind and belongs in the list anyway: on 20 July, Anthropic changed which Claude subscriptions include Claude Fable 5. Not a model. It moves the same routing variable the models move, which is the point of this piece.

The reflex is to ask which model is best. That question has no operational answer, because nobody runs a model. They run a set of routes. A team with agents in production already decided that some work goes to a fast cheap lane, some to a mid lane, some to the strongest thing available. Every one of those choices was made under specific prices, context limits, output caps, latency profiles, retention terms and access rules. A wave moves several of those inputs at once, in different directions, for different vendors. What it invalidates is not your model choice. It is the arithmetic under your model choice.

So the question that produces work is narrower: for each class of work you already run, what did this week change about the cost of that lane, the risk of that lane, and the cost of leaving it. Sometimes the answer is nothing, and nothing is a legitimate finding worth a dated line in a table.

Figure 03

Nine days, six releases

  1. 16–17 JulKimi K3sources conflict

    Moonshot's blog and Forbes say 17 July; OpenRouter and Simon Willison say 16 July. Shown as a range because the sources genuinely conflict.

  2. 16–21 JulQwen-Image-3.0sources conflict

    Blog metadata says 16 July, the API doc was last updated 20 July, press converges on 21 July. No primary source states a date in body text.

  3. 19 JulQwen3.8-Max-Previewconfirmed

    Previewed at WAIC Shanghai. Token Plan only, Beijing region only.

  4. 20 JulFable 5 subscription inclusionconfirmed

    Max, premium Team and premium seat-based Enterprise, capped at 50% of weekly limits. A commercial change, not a capability one.

  5. 21 JulGemini 3.6 Flashconfirmed

    Generally available. The only release in the wave with a benchmark table that could be extracted as text.

  6. 24 JulClaude Opus 5confirmed

    $5 / $25 per million — the same price as Opus 4.8, and half of Fable 5. Benchmarks published as charts only.

  7. 25 JulThis notetoday

    Publication date. Everything above is stamped to it.

  8. 27 JulKimi K3 weights, promisedpromised

    Moonshot committed to full weights by this date. Not published as of 25 July. A commitment, not an accomplishment.

  9. Qwen3.8-Max weights, promisedpromised

    Announced as coming “soon”. No date and no named licence.

Disputed dates are drawn as ranges rather than points, and forward-dated commitments are drawn as hollow markers. Both are findings about the releases, not gaps in the reporting.

  • A wave can change unit price, output cap, context limit, latency, access terms, retention terms, or region. Each is a separate routing input.
  • A model that is better on every axis and unavailable in your region changed nothing for you.
  • The unit of a routing decision is a work class, not a product name.
  • Re-pricing is triggered by an input moving, not by a launch happening.
  • 'No change' is a valid, dated, recordable outcome.

02 / The receipt on this piece

One model in this wave published a benchmark number we could check.

We publish the method with the conclusion. Every figure below comes from a vendor's own documentation, a vendor's own announcement, or an independent measurement page, each fetched directly on 25 July 2026. Where a page could not be fetched or a number was not published, we say so and we do not estimate.

The most useful finding is a negative one. Of the five models, exactly one published numeric benchmark results we could extract as text: Gemini 3.6 Flash, on its DeepMind model card. Anthropic's Opus 5 announcement presents every result as a chart, and we could not pull a single numeric score out of it. Moonshot's Kimi K3 post carries a section headed as a full benchmark table whose numbers are rendered client-side and appeared in none of our fetches. Alibaba published no benchmark at all, for either Qwen release. Four of the five models in this wave shipped no benchmark number we could extract from a fetched page. Any score we printed for those four would be a chart transcription or an aggregator's copy, and we print neither.

Second finding: the dates conflict. Kimi K3 is dated 17 July 2026 by Moonshot's own blog and by Forbes, and 16 July 2026 by OpenRouter and by Simon Willison. The likeliest explanation is an announcement in UTC+8 straddling a date boundary. We do not adjudicate; we write mid-July. Qwen-Image-3.0 is messier: press converges on 21 July, the official blog page metadata carries 16 July, and Alibaba Cloud's English API reference reports a last update of 20 July. No primary source states a release date in body text at all.

Third finding: two of the five were announced with open-weights intentions, and neither had public weights when we checked. Kimi K3's licence is stated on no page we could fetch, and with no public weights there is no public licence file to read either. Anthropic's Opus 5 system card exceeded our fetch size limit, so nothing in it is verified here, including the safety level that secondary outlets report. Everything in this piece is stamped 25 July 2026, and some of it will be stale within days.

Figure 01

What each vendor actually published

Verification quality of six model releases across five checkable attributes
ReleaseNumeric benchmarks extractablePrice publishedContext / output limitsLicence publishedDate unambiguous
Gemini 3.6 FlashYesYesYesn/a — closedYes
Claude Opus 5Charts onlyYesYesn/a — closedYes
Claude Fable 5Charts onlyYesYesn/a — closedYes
Kimi K3Table is JS-renderedYesYesNo16 vs 17 July
Qwen3.8-Max-PreviewNone publishedCredits onlyNoNoYes
Qwen-Image-3.0None publishedFree, limited timePixel bounds onlyNo16 vs 20 vs 21 July

Checked 25 July 2026. “n/a — closed” means no weights are distributed, so no model licence applies; terms of service still govern use. The column that matters is the first one: for five of six releases, no benchmark number could be extracted from a primary source at all.

  • Numeric benchmarks we could extract: Gemini 3.6 Flash only.
  • Charts with no extractable numbers: Claude Opus 5.
  • Benchmark table present but rendered client-side, not extractable: Kimi K3.
  • No benchmarks published at all: Qwen3.8-Max-Preview, Qwen-Image-3.0.
  • Licence not published: Kimi K3, Qwen3.8-Max-Preview, Qwen-Image-3.0.
  • Announcement date conflicting across sources: Kimi K3, Qwen-Image-3.0.

Fast lane

Extraction · formatting · deterministic checks

Balanced lane

Default implementation · routine research

Frontier lane

Architecture · ambiguity · security · synthesis

Human gate

Consequential action · approval · acceptance

03 / Four lanes

Define a work class by what its failure looks like, not by how hard it feels.

The extraction lane covers classification, extraction, normalization, formatting and schema conformance. Its failure mode is silent and well-formed: a wrong field, a dropped row, a plausible category on an implausible input. The output looks correct, so reading it is the wrong check. Assertions, fixtures and sampling are the right one, and that is why this lane tolerates a cheaper model than intuition suggests. The safety net is not the model. It is the check.

The implementation lane covers routine code, tests, migrations of a kind you have done before, documentation and small-diff review. Its failure mode is loud: it breaks CI or gets caught in review. We price this lane in review minutes rather than tokens, which means a model that produces slightly worse code and much shorter diffs can win outright.

The frontier lane covers architecture, ambiguous debugging, unfamiliar migrations, security-sensitive synthesis and any decision where being wrong costs more than every token you will ever spend on it. Its failure mode is the dangerous one: a confident, coherent, internally consistent answer that is wrong in a way the reviewer lacks the context to detect. This lane is priced by consequence. Token cost is not the quantity that decides it.

The multimodal lane is really two lanes. Understanding takes images, audio, video or documents in and returns text. Generation returns an artifact. They have different vendors, different failure modes, different cost structures and, this month, very different licensing clarity. Treating them as one lane is a routing mistake, and this wave makes the split easy to see.

  • Extraction: silent structural failure. Cheap lane plus hard assertions.
  • Implementation: loud failure caught by CI and review. Priced in review minutes.
  • Frontier: confident wrong answers. Priced by consequence, not tokens.
  • Multimodal understanding: routes like extraction with a larger bill.
  • Multimodal generation: routes like a supply-chain decision, because licence and retention dominate.

04 / The cheap lane, re-priced

Gemini 3.6 Flash re-priced the cheap lane. The output cap decides who can use it.

Google announced Gemini 3.6 Flash on 21 July 2026; its docs state it is generally available and ready for production, under the API model ID gemini-3.6-flash. Standard pricing is $1.50 per million input tokens and $7.50 per million output tokens, with output including thinking tokens. Batch is $0.75 and $3.75, flex identical to batch, priority $2.70 and $13.50. Context caching costs $0.15 per million input tokens plus $1.00 per million tokens per hour of storage, so the storage term, not the read rate, is the number to model. The context window is 1M input with a 64K output limit. The knowledge cutoff is March 2026, with the model card noting some domains limited to January 2025. Inputs may be text, images, audio or video, with PDFs listed on the DeepMind Flash page. Output is text only.

The four price tiers are the actual product: work that can wait should not pay the synchronous rate, and batch is half of standard. Google also claims 3.6 Flash reduces output token usage by 17% against 3.5 Flash, per the Artificial Analysis Index. Read that carefully, because it is the number most likely to be misquoted this month: it is a reduction in tokens consumed, not a price cut. For output-billed work it is arguably the more interesting number, since consumption is the term you cannot negotiate. Secondary reports also describe a cut to the output price, but we could not verify the prior figure on any Google primary source, so we neither print it nor lean on it. Google further cites Datacurve reporting up to 65% output token reduction on the DeepSWE benchmark, and partner figures from Harvey (12% faster task completion) and JetBrains (10 to 20% coding performance gains).

The 64K output cap is the specification that decides whether a job belongs in this lane. For extraction it almost never binds. For any job whose artifact exceeds 64K tokens it is a hard wall, and no amount of intelligence compensates. For contrast: Claude Opus 5 emits up to 128k tokens synchronously and up to 300k on the Message Batches API behind a beta header, and Kimi K3 defaults to 131,072 output tokens and can be configured up to 1,048,576. Three 1M-context models whose output ceilings differ by a factor of sixteen. Context window is the number everyone quotes. Output cap is the number that routes work.

Two independent measurements from Artificial Analysis point in opposite directions and both matter. It measured 237.1 output tokens per second, ranked first of 190 models, and a time to first token of 14.73 seconds, which it describes as somewhat higher than average. A person watching a cursor experiences the first token; a queue draining output experiences the rate. The same model ranks first on one measurement and below average on the other. Artificial Analysis scores it 50 on its Intelligence Index v4.1, ranked 26th of 190 against a median of 32.

  • SWE-Bench Pro: 58.7%, against Gemini 3.5 Flash at 55.1%.
  • DeepSWE v1.1: 49%, against 3.5 Flash at 37%.
  • Terminal-bench 2.1: 78.0%.
  • MLE-Bench: 63.9%, against 3.5 Flash at 49.7%.
  • OSWorld-Verified: 83.0%, against 3.5 Flash at 78.4%.
  • GDM-MRCR v2 long context: 91.8% at 128k, and 54.0% at 1M pointwise. A window that retrieves at 54% is not a 1M window for correctness-critical work.

05 / The frontier lane, re-priced

Opus 5 is not a price cut in the Opus line. It is a claim about the frontier lane a canary can check.

Anthropic announced Claude Opus 5 on 24 July 2026 at $5 per million input tokens and $25 per million output tokens. That is exactly what Claude Opus 4.8 cost, and exactly half of Claude Fable 5 at $10 and $50. Anthropic's framing is that Opus 5 comes close to the frontier intelligence of Fable 5 at half the price, and that it delivers greatly improved performance for the same cost as its predecessor. Both statements describe the same price list. If your frontier lane runs on Fable 5, the vendor is now arguing that the lane costs half. Whether it does on your workload is what a canary is for, because a spec sheet does not carry pass rates, retry burden, or review minutes.

There are no numeric benchmark scores to check. The announcement presents everything as charts, and the named evaluations are Frontier-Bench v0.1, CursorBench 3.2, the AA Coding Agent Index, ARC-AGI 3, GDPval-AA v2, OSWorld 2.0, HLE, Zapier AutomationBench and DeepSearchQA. What we could verify is Anthropic's own wording about relative position, with no absolute numbers behind it: on Frontier-Bench v0.1 it surpasses all other models and more than doubles Opus 4.8's performance at a lower cost per task; on CursorBench 3.2 at max effort it is within 0.5% of Fable 5's peak score at half the cost; on ARC-AGI 3 its score is three times the next-best model; on Zapier AutomationBench its pass rate is around 1.5x the next-best model for the same cost; on OSWorld 2.0 it surpasses Fable 5's best result at just over a third of the cost. Four of those five claims are cost-normalized. The vendor is arguing on the same axis this piece argues on, which is interesting, and unfalsifiable from outside without the numbers, which is the problem.

The effort parameter is the real routing dial, and it means the price list is not the price. Opus 5 supports five levels: low, medium, high, xhigh and max, with high as the API default on both the Claude API and Claude Code, set through output_config in the Messages request. Thinking cannot be disabled at xhigh or max; a request that tries returns a 400. Anthropic also notes that on Opus 5, changing effort does not reliably shorten the visible response, and that you should prompt for length instead. Effort moves token consumption, and token consumption is what you are billed for. A cost model that treats $5 and $25 as constants is modeling the list, not the bill.

Two line items on the rest of the surface move the bill by multiples; the others are small constants worth knowing. Caching first: cache hits and refreshes cost $0.50 per MTok against $5.00 fresh, and a five-minute cache write costs $6.25, a premium a single reuse inside the window more than repays. The one-hour write costs $10.00 and needs at least two reads against the same prefix to come out ahead. Effort is the second, covered above. The small constants: batch at $2.50 and $12.50; fast mode at $10 and $50 for around 2.5 times default speed, in research preview, first-party Claude API only, not available on Claude Platform on AWS, partner clouds, or the Batch API; the full 1M window billed at standard rates with no long-context surcharge on Claude 4.6 and later; tool use adding 286 input tokens of system prompt overhead for tool_choice auto or none and 406 for any or tool, with the Bash tool adding a further 325; and Claude Managed Agents at $0.08 per session-hour on top of tokens. None of that is exciting. All of it shows up on the invoice.

  • Opus 5: $5 / $25 per MTok, 1M context, 128k max output synchronous, 300k on Batches behind a beta header, May 2026 training cutoff.
  • Fable 5: $10 / $50 per MTok, 1M context, 128k max output, January 2026 cutoff.
  • Same price as Opus 4.8. Half the price of Fable 5. Not a price cut in the Opus line.
  • Effort default is high; thinking cannot be disabled at xhigh or max.
  • Fast mode: 2x price for roughly 2.5x speed, research preview, first-party API only.
  • Opus 5 remains behind Claude Mythos 5 on cybersecurity tasks and biology research, per Anthropic.

06 / Access is a routing input

The Fable 5 change is not a capability change. It is the access boundary moving.

From 20 July 2026, Claude Fable 5 is included at no extra cost on Max, on premium Team seats, and on premium seat-based Enterprise seats, capped at 50% of weekly usage limits. That cap is not additive: other models draw from the same weekly pool, and the total is still the weekly limit. Pro and Team Standard seats are not included and pay through usage credits after a one-time credit. Anthropic's help centre says only that a one-time credit exists; the @claudeai account, as quoted by Simon Willison, gives the figure as $100. The prior free-access promotion ended 19 July 2026 at 11:59:59 PM PT. One wording matters here: the help-centre article never uses the word permanent. It states a change effective 20 July with no end date, and the 'permanent' in the headlines is the journalists' word, not a quote. An absent end date is not a commitment, and the difference between those two readings is the whole question of whether you can build on this.

The instability is itself the data. Anthropic revised Fable 5's plan access at least four times between 9 June and 20 July 2026, and its own stated reason is that demand proved hard to predict, so access was extended in stages as capacity was secured. Separately, Fable 5 was forced offline on 12 June 2026 by a US export-control directive and returned on 1 July. Independent reporting describes the practical effect of the July change as 50% of already-reduced limits, following a roughly one-third cut to regular usage when the bonus phase ended. None of that says anything about the model's capability, and all of it is a routing input. A lane whose commercial terms moved four times in six weeks, and which a government directive took offline once, should not be the only lane a team is able to run.

Two constraints in this release beat capability outright. Fable 5 is designated a Covered Model with 30-day data retention and is not available under a zero-data-retention agreement. For a regulated workload that is dispositive, and no benchmark can argue with it. The second is quieter and costs teams real money: Fable 5 uses the tokenizer introduced with Claude Opus 4.7, and the same text produces roughly 30% more tokens than on models released before 4.7. Dollars per million tokens is therefore not a comparable unit across vendors, or even across generations of one vendor. Price your lanes on your own corpus, in tokens you counted yourself, or you are comparing two different rulers.

  • Included: Max, premium Team seats, premium seat-based Enterprise seats, up to 50% of the weekly pool.
  • Not included: Pro, Team Standard. Usage credits after a one-time credit.
  • The help-centre article does not call the change permanent. It states a start date and no end date.
  • Fable 5 is a Covered Model: 30-day retention, no zero-data-retention option.
  • Fable 5's tokenizer produces roughly 30% more tokens for the same text than pre-4.7 models.
  • Plan list prices as published: Pro $17/mo annual or $20/mo monthly; Max from $100/mo; Team Standard $20/seat annual or $25/seat monthly; Team Premium $100/seat annual or $125/seat monthly; Enterprise $20/seat.

07 / Open weights as posture

Two open-weights commitments, zero public weights.

Kimi K3 is real, hosted, and technically the most interesting thing in the wave. Moonshot's blog describes a 2.8 trillion parameter model using a sparse mixture of experts that activates 16 of 896 experts per token, an attention design named Kimi Delta Attention with attention residuals, and quantization-aware training with MXFP4 weights and MXFP8 activations. vLLM's preview write-up corroborates the expert configuration and the MXFP4 weights, and describes the attention as KDA-dominant linear attention with periodic full-attention layers. Inputs are text, images and video files per the API docs. Context is 1,048,576 tokens. Pricing is $3.00 per million input tokens on a cache miss, $0.30 on a cache hit, and $15.00 per million output, flat, with no tiering by context length. Reasoning is always on, with effort levels low, high and max, defaulting to max. API access unlocks after a minimum $1 top-up. Moonshot positions it for long-horizon coding and end-to-end knowledge work.

What is missing is the part that matters for resilience. As of 25 July 2026 the moonshotai organisation on Hugging Face lists no K3 repository; the newest entries are K2.7-Code and older. Moonshot's blog commits to releasing full weights by 27 July 2026, and vLLM restated that as still pending on 22 July. Third-party repositories matching the name exist and are not Moonshot's. No page we fetched states a licence, and with no public weights there is no public licence file to read. vLLM's support is preview-stage and requires large-scale expert parallelism. None of this stops a team calling the hosted API today; the model ID, prices and limits are published and the lane is usable. What it stops is the thing open weights are for: running it, pinning it, and no longer caring what the vendor does next. That part is a date on a calendar, two days after this piece publishes.

Qwen3.8-Max-Preview is thinner still. Alibaba's Qoder documentation states 2.4 trillion parameters. It was previewed on 19 July 2026 at the World Artificial Intelligence Conference in Shanghai. Access runs through Alibaba's Token Plan subscription, Qoder and QoderWork, and Token Plan is available only in the China (Beijing) region. There is no published per-token price; billing is credits-based, with the preview discounting the billing coefficient from 0.5x to 0.05x, and to 0.01x between 22:00 and 08:00 UTC+8. Alibaba published no model card, no benchmark table, no activated-parameter count and no licence, and we confirmed the absence directly: the Model Studio specification tables contain no row for this model. The only performance claim Alibaba made is a single post from the Qwen team describing it as second only to Fable 5, with nothing published behind it. Weights are promised soon, with no date and no named licence. One trap worth flagging: a context-window figure for this model circulates widely and appears on no Alibaba page we fetched; the 1M figure in those tables belongs to Qwen3.7-Max, a different model. We state no context window for Qwen3.8-Max-Preview, because we do not have one.

Two vendors announced very large models with open-weights intentions inside one week, and neither has shipped public weights. Provider resilience is not a second API key, and it is certainly not a promised weight drop. It is three concrete things: a tool contract portable across providers, a pinned replay set that defines success independently of any model, and a second lane you have actually run under production load at least once. Build all three as if the second lane will be needed. Then treat every announced weights release as a diary entry to verify, not a capability you already possess.

  • Kimi K3: 2.8T total parameters, 16 of 896 experts active, 1,048,576-token context.
  • Kimi K3 pricing: $3.00 input cache miss, $0.30 cache hit, $15.00 output, per million tokens, flat.
  • Kimi K3 weights: promised by 27 July 2026, not public as of 25 July 2026. Licence unstated.
  • Kimi K3 is callable today as a hosted lane. It is not pinnable, and that is the difference.
  • Qwen3.8-Max-Preview: 2.4T parameters claimed, preview only, Token Plan only, Beijing region only.
  • Qwen3.8-Max-Preview: no benchmark, no licence, no per-token price, no verified context window.
  • Resilience is a portable tool contract, a pinned replay set, and a second lane you have actually run.

08 / Multimodal is two lanes

Understanding got broader. Generation got a licence problem.

On the understanding side, this wave was generous. Gemini 3.6 Flash accepts text, images, audio and video, with PDFs listed on the DeepMind Flash page, and returns text only. Kimi K3 accepts text, images and video files, with images supplied as base64 or an ms:// file identifier; its docs state inputs and make no statement about output modality beyond text responses being the default. Qwen3.8-Max-Preview is reported to handle text, images, video and document understanding, but that comes from secondary coverage only and we could not confirm it on any Alibaba-owned page. Understanding routes like the extraction lane with a larger bill: emit a structured object, assert on the object, and never let a reviewer's impression of the prose stand in for a check.

Generation is a different discipline. Qwen-Image-3.0 is exposed on Alibaba Cloud Model Studio as qwen-image-3.0-pro, handling text-to-image and image-to-image in one model, with image-to-image taking one to three reference images plus a single instruction. It is in limited preview and requires an approved application through the Model Gallery. It serves from two regions, China (Beijing) and Singapore, with separate keys and endpoints that cannot be used interchangeably. Verified API constraints: total output pixels between 512x512 and 2048x2048, PNG output, n between 1 and 6 images with a default of 1, prompt_extend defaulting to true, watermark defaulting to false, seed in the range 0 to 2147483647. Input images may be JPG, JPEG, PNG, BMP, TIFF, WEBP or GIF, recommended between 384 and 2048 pixels per side, up to 10MB. Requests are single-turn with exactly one user message.

The vendor claims are the interesting part and none of them is checkable. Qwen's blog states support for prompt input up to 4.5k tokens, precise rendering of text as small as 10px, native rendering of 12 languages, and over 100 artistic styles; it offers an example of a three-by-three grid of infographics generated in a single pass from a 3.7k-token prompt rather than stitched together. There is no benchmark table, no parameter count, no architecture description, no technical report and no licence statement on the blog or in the API reference. The price is listed as free for a limited time in both regional rate tables, which is not a price. For scale, the same tables put the prior generation, qwen-image-2.0-pro, at ¥0.5 per image in Beijing.

Two of those facts are pipeline design rather than trivia. Generated image URLs expire after 24 hours, so any system that retains images must persist them on receipt; a pipeline that stores the URL has a 24-hour fault built into it that will not surface until someone opens an old record. And no licence is published for the model or for commercial use of its outputs. Our policy is blunt on this: we do not route client deliverables through a generation lane whose usage terms we cannot produce in writing, and nothing published for Qwen-Image-3.0 clears that bar yet. The Qwen image line's last open weights were Qwen-Image-2512, Apache 2.0, dated 31 December 2025. Version 3.0 is closed, invite-gated, and unlicensed in public. That is a supply-chain decision before it is an aesthetic one.

  • Understanding in: Gemini 3.6 Flash (text, image, audio, video, PDF), Kimi K3 (text, image, video).
  • Understanding out: text only for Gemini 3.6 Flash and the Claude models. Kimi K3 and Qwen3.8-Max-Preview publish no output-modality statement.
  • Generation: Qwen-Image-3.0 via qwen-image-3.0-pro, limited preview, application required.
  • Output ceiling 2048x2048 total pixels, PNG. Do not expect native 4K.
  • Generated URLs expire after 24 hours. Persist on receipt.
  • No licence published for Qwen-Image-3.0 or its outputs. Under our policy that keeps it out of client work, whatever its quality.

09 / The policy

A routing policy you can adopt this week and measure next month.

Write the policy as a table of work classes, not model names. For each class, record six fields: the default lane, the escalation trigger, the escalation lane, the hard constraints (retention, region, output cap, licence), the fixture set that proves the lane works, and the date the lane was last re-priced. A policy that lists model names without work classes is a preference. A policy that lists work classes without dates is a memory. The dates are what let you answer, in six weeks, whether a decision was made under conditions that still hold.

Then apply the wave. Extraction and formatting default to a cheap lane, and if that lane is Gemini 3.6 Flash, route asynchronous volume to batch at $0.75 and $3.75 rather than standard at $1.50 and $7.50, and check that no job in the class needs more than 64K output. Routine implementation stays wherever your review-minute cost is lowest; nothing verified in this wave forces a move, and a lane that produces shorter diffs beats a lane that scores better. Frontier work routes to the strongest lane you can afford per task, and if that lane was Fable 5, run a canary on Opus 5 at $5 and $25 with effort at xhigh before treating the halved list price as a halved cost. Multimodal understanding joins the extraction lane with harder assertions. Multimodal generation stays out of client work until usage terms exist in writing.

Measure five things. Cost per accepted task, not cost per million tokens, because effort levels and tokenizers make per-token comparisons meaningless across vendors. Escalation rate. Escalation win rate, meaning how often the frontier lane actually changed the answer. Review minutes per accepted task. And the date of your last failover rehearsal. If escalation rate is high and escalation win rate is low, the lane is fine and the trigger is miscalibrated; fix the trigger, not the lane. What this piece cannot hand you is thresholds. Those come out of your telemetry, and any consultant who quotes one without seeing your telemetry is guessing.

The standing rule is that you re-price when an input moves, not when a model launches. Price, context limit, output cap, access terms, retention terms, region: any of those moving opens the table. The rule needs one honest amendment: a capability claim at unchanged list price is not a re-route trigger, but it is a canary trigger, because cost per accepted task can fall while the price list holds still, through fewer retries and shorter reviews. The re-route still waits for the canary's numbers, and the canary's arithmetic must include the cost of switching itself: integration work, prompt portability, cache loss, and reviewers learning a new failure profile. By our reading of this week: one release re-priced the cheap lane, one re-priced the frontier lane pending canary, one moved an access boundary without touching a model, one added a hosted lane whose weights and licence are promises with a date attached, and two we cannot act on at all, one gated to a region and subscription we do not hold and one invite-gated with no licence. The first of those promises comes due on 27 July. Write down what moved, what you did about it, and what you could not verify, and date all three.

  • Route extraction to the cheap lane with hard assertions. Use batch pricing for asynchronous volume.
  • Check the output cap before the context window. 64K, 128K, 300K and 1,048,576 are all in this wave.
  • Route implementation by review minutes, not by benchmark position.
  • Route frontier work by consequence. Canary a cheaper frontier lane before adopting it.
  • A capability claim at constant price is a canary trigger, not a re-route trigger.
  • Keep generation out of client deliverables until usage terms are published.
  • Re-price on input movement. Record 'no change' with a date when nothing moved.
  • Rehearse failover to the second lane on a schedule, and record the date it last succeeded.

10 / Production missions

Prompts that specify a whole mission.

These are not idea starters. Each is an operating contract for GPT-5.6 or Fable 5: outcome, boundaries, architecture, model policy, phases, acceptance tests, evidence, release, and stop conditions.

Mission 01 · Model routing

Release-wave re-pricing audit

A dated routing table that states which lanes a release wave actually changed, which it did not, and what could not be verified.

Claude Opus 5 · effort high · pinned snapshot

Full operating contract

MISSION: RE-PRICE EVERY ROUTING DECISION AFTER A MODEL RELEASE WAVE

Act as the routing owner for a team running agents in production. A set of model releases has landed. Your job is not to summarise them. Your job is to determine, for each work class we already run, whether any routing input moved, and to leave a dated record that a future engineer can audit.

INPUTS I WILL PROVIDE
- Our current routing table: work classes, default lane, escalation trigger, escalation lane, hard constraints.
- 30 to 90 days of production telemetry: tokens in and out per class, effort levels used, cache hit ratio, latency percentiles, escalation events, rejected outputs.
- The releases in scope, with vendor documentation URLs.
- Our compliance constraints: data retention, region, licence requirements, approved providers.

OPERATING CONTRACT
1. Treat vendor announcements as claims. Record the URL and the fetch date for every number you use.
2. If a number is not published, write UNVERIFIED and say what page you checked. Never estimate a price, a context limit, an output cap, or a benchmark score.
3. Distinguish list price from realised cost. Effort parameters, thinking tokens, tool-definition overhead, cache write premiums and tokenizer differences all move realised cost.
4. Never compare dollars per million tokens across vendors without re-tokenizing a sample of our own corpus. State the token ratio you measured.
5. A capability claim with no change to price, limits, access or risk is not a re-route trigger by itself. It is a canary trigger, because cost per accepted task can move at constant list price. The re-route waits for the canary's numbers.

PRODUCE, PER WORK CLASS
- The six routing inputs and whether each moved: unit price, context limit, output cap, latency profile, access terms, retention and region.
- Realised cost per accepted task under the current lane, computed from our telemetry rather than from the price list.
- Projected realised cost under each candidate lane, with the assumptions stated as a list.
- A hard-constraint check that can veto a cheaper lane outright: retention terms, region availability, output cap, published licence or usage terms.
- A verdict of one of: no change, canary recommended, migrate, or blocked pending verification.
- The date, and the specific fact that would force a re-check.

ADVERSARIAL PASS
- Identify every place where a cheaper lane would fail silently rather than loudly, and say what assertion would catch it.
- Identify every lane where we currently have no tested alternative, and name the risk in one sentence.
- List any claim in the vendor material that is cost-normalized but unfalsifiable without absolute numbers.

DELIVERABLE
A single dated routing table plus an UNVERIFIED appendix listing every number we wanted and could not get, with the page we checked. Do not recommend a migration for any class where the appendix contains a number that migration decision depends on.

Mission 02 · Cost engineering

Cheap-lane migration canary

A measured move of one work class to a cheaper lane, with an escalation trigger that either pays for itself on evidence or reverts.

Candidate cheap lane under test · stronger lane as adjudicator

Full operating contract

MISSION: MOVE ONE WORK CLASS TO A CHEAPER LANE WITHOUT LOSING THE PLOT

Act as the engineer responsible for a single work class. We believe a cheaper lane can carry it. Prove or disprove that with evidence, and leave behind a trigger that routes the residual cases upward.

INPUTS I WILL PROVIDE
- The work class definition, its current lane, and its acceptance criteria.
- 100 to 300 historical cases with known-good outcomes, including edge cases, contradictory inputs, missing data, and at least ten adversarial or prompt-injection attempts.
- Current realised cost, latency percentiles, and review minutes per accepted task.
- The candidate lane's published price, context limit, output cap, and constraints.

OPERATING CONTRACT
1. Pin model identifiers on both lanes. Never compare a rolling alias against a frozen baseline.
2. Check the output cap against the largest artifact in the fixture set before running anything. If any case exceeds it, the migration is dead and you should say so in the first paragraph.
3. Success is defined by the fixtures, not by the model. Do not change the acceptance criteria to fit the cheaper lane.
4. Measure cost per accepted task, including retries, escalations, and the review minutes a human spends. Not tokens.
5. Run the asynchronous portion through batch pricing where the vendor offers it, and report the split.
6. Include the switching cost itself in the arithmetic: integration work, prompt portability, cache loss, and reviewer retraining.

BUILD
- A parallel run harness that sends every fixture to both lanes and records output, tokens, latency to first token, latency to completion, and cost.
- A deterministic differ that classifies each disagreement as: cheaper lane wrong, stronger lane wrong, both acceptable, or ambiguous requiring human adjudication.
- An escalation trigger expressed as a rule over observable signals, not over model confidence. Candidate signals: schema violation, low field coverage, contradiction between extracted fields, input length or structure outside the fixture distribution, retry count.
- A shadow-mode deployment that routes production traffic to the cheap lane, escalates on the trigger, and logs both results without changing user-visible behaviour.

ACCEPTANCE
- Escalation rate and escalation win rate are both reported. A high escalation rate with a low win rate means the trigger is wrong, not the lane.
- No adversarial fixture produces an accepted output on either lane.
- Realised cost per accepted task is lower on the new configuration including escalations and switching cost, or the migration is rejected.
- A single configuration change reverts to the previous lane, and that revert has been executed once in staging.

DELIVERABLE
The parallel-run results table, the trigger rule, the shadow-mode logs, the revert procedure with its rehearsal date, and an explicit recommendation. If the recommendation is to stay, say so plainly and state what would change the answer.

Mission 03 · Provider resilience

Second-lane failover drill

Evidence that a named second lane can carry a work class today, rather than a second API key nobody has ever exercised.

Primary lane and declared fallback lane, both pinned

Full operating contract

MISSION: PROVE THE SECOND LANE WORKS BEFORE YOU NEED IT

Act as the reliability owner for a model-dependent system. We have a declared fallback provider. Determine whether it is real. A configuration entry is not a fallback; a lane that has carried production traffic under load is.

CONTEXT THAT MOTIVATES THIS
Access terms are a routing input and they move independently of capability. In one six-week window in mid-2026, one frontier model had its subscription terms revised at least four times and was separately taken offline for weeks by an export-control directive. In the same month, two vendors announced open-weights models and had shipped no public weights at the time of checking. Plan for the access boundary moving, not for the model getting worse.

INPUTS I WILL PROVIDE
- The work classes in scope and their current primary lanes.
- The declared fallback for each, and the tool contract each lane is called through.
- The pinned replay set that defines success for each class.
- Compliance constraints: retention, region, licence, approved providers.

OPERATING CONTRACT
1. Never treat a promised open-weights release as an available fallback. Only a lane you can call today counts.
2. Check the fallback against the hard constraints first: retention terms, region availability, output cap, published licence or usage terms. A fallback that fails a compliance check is not a fallback.
3. The tool contract must be identical across lanes. If the fallback requires a different tool schema, that difference is the actual work item.
4. Record realised cost on the fallback lane. A fallback you cannot afford for a week is a fallback you will not use.

RUN THE DRILL
- Force a hard failover of one work class to the fallback lane for a bounded window on real traffic.
- Measure pass rate against the pinned replay set, realised cost per accepted task, latency to first token and to completion, and review minutes.
- Inject the failure modes that actually occur: provider 429s, elevated latency, a refusal-style stop reason returned with a success status code, a region becoming unavailable, and an account or plan-level access change.
- Fail back, and verify that in-flight work either completed or was cleanly retried with no duplicated external side effects.

DELIVERABLE
A one-page drill record: date, work class, fallback lane, pass rate delta, cost delta, latency delta, failures observed, defects filed, and the next scheduled drill date. State plainly which work classes have no viable second lane today and what it would take to create one. Do not report success for any class where the drill did not actually run.

11 / En español

Resumen en español

Entre el 16 y el 24 de julio de 2026 salieron cinco modelos y se movió una frontera de acceso: Kimi K3, Gemini 3.6 Flash, Claude Opus 5, Qwen3.8-Max-Preview, Qwen-Image-3.0, y el cambio del 20 de julio en qué suscripciones de Claude incluyen Claude Fable 5. La reacción natural es preguntar cuál es el mejor. Esa pregunta no tiene respuesta operativa. Nadie opera un modelo: se operan rutas, y una ola de lanzamientos no te entrega un ganador, te invalida la aritmética con la que decidiste esas rutas.

Conviene rutear por clase de trabajo. Extracción, clasificación y formato: la falla es silenciosa y bien formada, así que el carril puede ser barato siempre que las aserciones sean duras. Implementación rutinaria: la falla la atrapan CI y la revisión, y el costo real se mide en minutos de revisión. Frontera (arquitectura, depuración ambigua, migraciones desconocidas): la falla es una respuesta coherente y equivocada que el revisor no tiene contexto para detectar, así que se cotiza por consecuencia, no por tokens. Multimodal: en realidad son dos carriles, comprensión y generación, con proveedores, fallas y licencias distintas.

Los números verificados que sí mueven decisiones: Gemini 3.6 Flash a $1.50 y $7.50 por millón de tokens, con lote a $0.75 y $3.75, ventana de 1M de entrada y tope de 64K de salida. Claude Opus 5 a $5 y $25, el mismo precio que Opus 4.8 y la mitad de Fable 5 ($10 y $50); no es una rebaja en la línea Opus, es la afirmación del proveedor de que el carril de frontera cuesta la mitad, y eso se comprueba con un canario, no con el anuncio. Kimi K3 a $3.00 de entrada sin caché, $0.30 con caché y $15.00 de salida, con ventana de 1,048,576 tokens. El tope de salida importa más que la ventana de contexto para decidir rutas.

También hay que decir lo que no se pudo verificar, porque es igual de útil. De los cinco modelos, sólo Gemini 3.6 Flash publicó benchmarks extraíbles como texto. Anthropic presentó los suyos como gráficas. La tabla de Kimi K3 se renderiza por JavaScript y no se pudo extraer. Alibaba no publicó ningún benchmark para sus dos lanzamientos. La fecha de anuncio de Kimi K3 se contradice entre fuentes (16 contra 17 de julio) y su licencia no aparece en ninguna página. Los pesos abiertos de K3 estaban prometidos para el 27 de julio de 2026 y al 25 de julio no existían en público. Los de Qwen3.8-Max se prometieron sin fecha y sin licencia.

De ahí sale la política: la resiliencia de proveedor no es una segunda API key ni una promesa de pesos abiertos. Son tres cosas concretas: un contrato de herramientas portable, un conjunto de casos fijados que define el éxito sin depender del modelo, y un segundo carril que ya corriste bajo carga real por lo menos una vez. Y la regla permanente: se recotiza cuando se mueve un insumo (precio, tope de salida, términos de acceso, retención, región), no cuando alguien lanza un modelo; una promesa de capacidad a precio constante amerita canario, no migración.

  • Rutea por clase de trabajo, no por nombre de modelo. La clase se define por cómo falla.
  • Revisa el tope de salida antes que la ventana de contexto: 64K, 128K, 300K y 1,048,576 conviven en esta ola.
  • El precio de lista no es el precio: el parámetro de esfuerzo y el tokenizador mueven el consumo real.
  • Fable 5 es Covered Model con retención de 30 días y sin opción de retención cero. Eso decide antes que cualquier benchmark.
  • No cites un score que no puedas extraer de una fuente primaria. Cuatro de los cinco modelos de esta ola no publicaron ninguno en forma extraíble.
  • Mide costo por tarea aceptada, tasa de escalamiento, tasa de acierto del escalamiento, minutos de revisión y fecha del último simulacro de failover.
  • Registra 'sin cambio' con fecha. Es un resultado válido y mucho más barato que una migración.

12 / Source trail

Primary, fresh, inspectable.

Product details move quickly. These sources were checked for this issue on July 25, 2026. Re-verify before changing production policy.

  1. 01Kimi K3Moonshot AIPrimary launch post. Source for 2.8T parameters, 16-of-896 expert activation, MXFP4/MXFP8 training, native vision, pricing, and the commitment to publish full weights by 27 July 2026. Its benchmark table is JavaScript-rendered and could not be extracted.
  2. 02Kimi K3 quickstartMoonshot AIPrimary API reference. Source for the kimi-k3 model ID, always-on reasoning with low/high/max effort, accepted input modalities, output token defaults and ceiling, flat pricing structure, and the $1 minimum top-up.
  3. 03Kimi K3 pricingMoonshot AIPrimary source for the 1,048,576-token context window and the cache-miss, cache-hit and output rates.
  4. 04moonshotai on Hugging FaceHugging FaceChecked 25 July 2026. No Kimi K3 repository present; newest entries are K2.7-Code and earlier. This is the evidence that the weights were still unpublished.
  5. 05Kimi K3 preview supportvLLMDated 22 July 2026. Corroborates the expert configuration and MXFP4 weights, describes serving as requiring large-scale expert parallelism, and restates the weights release as still pending.
  6. 06Kimi K3Simon WillisonIndependent write-up dated 16 July 2026, one of the two sources that conflict with the vendor's 17 July date.
  7. 07Gemini 3.6 Flash model cardGoogle DeepMindThe only fully extractable numeric benchmark table in this wave. Source for context and output limits, knowledge cutoff, modalities, and every benchmark figure quoted here.
  8. 08Gemini API pricingGooglePrimary source for standard, batch, flex and priority rates, context caching costs, free tier, and grounding pricing.
  9. 09Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash CyberGooglePrimary announcement, 21 July 2026. Source for the 17% output-token reduction claim and the Datacurve figure.
  10. 10Gemini 3.6 FlashArtificial AnalysisIndependent measurement. Source for the Intelligence Index v4.1 score and rank, 237.1 output tokens per second, and the 14.73 second time to first token.
  11. 11Introducing Claude Opus 5AnthropicPrimary announcement, 24 July 2026. Source for pricing, positioning against Fable 5, fast mode, and the relative benchmark wording. All benchmark results are presented as charts; no numeric score was extractable.
  12. 12Model overviewAnthropicPrimary source for context windows, output caps, training cutoffs, model IDs, the Opus 4.7 tokenizer note, and the legacy status of Opus 4.8.
  13. 13PricingAnthropicPrimary source for cache write and hit rates, batch rates, fast mode pricing and availability limits, tool-definition token overhead, and Managed Agents session-hour billing.
  14. 14EffortAnthropicPrimary source for the five effort levels, the high default, the 400 error when disabling thinking at xhigh or max, and the note that effort does not reliably shorten visible output length.
  15. 15Claude Fable 5 on your planAnthropicPrimary source for the 20 July 2026 inclusion terms, the non-additive 50% cap, the exclusion of Pro and Team Standard, and the end of the free promotion. The article itself never uses the word 'permanent'.
  16. 16Introducing Claude Fable 5 and Claude Mythos 5AnthropicPrimary source for Fable 5 pricing, the 9 June 2026 availability date, always-on adaptive thinking, the refusal stop reason, and the Covered Model 30-day retention designation.
  17. 17Redeploying Fable 5AnthropicPrimary record of the 12 June 2026 export-control shutdown and the 1 July 2026 return to service.
  18. 18The @claudeai statement on Fable 5 plan inclusionSimon WillisonSecondary source carrying the @claudeai statement verbatim, including the $100 one-time credit figure that Anthropic's own help centre omits. The original post returned HTTP 402 and could not be fetched directly. The 'permanent' framing in this write-up is the author's, not a quoted Anthropic word.
  19. 19Anthropic slashes Claude Fable 5 limits in Max and Team PremiumThe DecoderIndependent account of the practical effect: Fable 5 capped at 50% of limits that had already been cut by roughly a third when the bonus phase ended.
  20. 20Qwen3.8-Max-Preview campaignQoder (Alibaba)The only Alibaba-owned page stating 2.4 trillion parameters. Source for the 19 July 2026 start date, eligibility, and the 0.5x to 0.05x and 0.01x billing coefficients.
  21. 21Token Plan overviewAlibaba Cloud Model StudioPrimary confirmation of the qwen3.8-max-preview model ID, its preview status, the China (Beijing) region restriction, and the Token Plan subscription tiers.
  22. 22Text generation modelsAlibaba Cloud Model StudioChecked directly. Contains no specification row for qwen3.8-max-preview, which is why this piece states no context window for that model. The 1M figure in these tables belongs to Qwen3.7-Max.
  23. 23Qwen-Image-3.0: Rich Content, Authentic Details, Deep KnowledgeQwenPrimary announcement. Source for the 4.5k-token prompt claim, 10px text rendering, 12 languages, and the single-pass infographic grid. Contains no benchmark table, parameter count, architecture detail, or licence statement.
  24. 24Qwen image generation and editing API referenceAlibaba Cloud Model StudioPrimary API reference for qwen-image-3.0-pro. Source for the limited-preview gating, the two non-interchangeable regional endpoints, the pixel bounds, and the single-turn request shape.
  25. 25Model Studio billingAlibaba CloudPrimary source for qwen-image-3.0-pro being listed as free for a limited time in both regions, and for the prior generation's ¥0.5 per image reference point.
  26. 26Qwen-Image on GitHubQwenLMChecked 25 July 2026. No 3.0 entry; the open-weight releases stop at Qwen-Image-2512 under Apache 2.0, dated 31 December 2025.
  27. 27Google releases three new Gemini models, but no 3.5 ProTechCrunchIndependent coverage of the 21 July 2026 release, the absence of a Gemini 3.5 Pro update, and the limited-pilot status of Gemini 3.5 Flash Cyber.
  28. 28Anthropic launches Opus 5TechCrunchIndependent confirmation of the 24 July 2026 launch date and its roughly two-month gap after Opus 4.8.
  29. 29Claude Opus 5 System CardAnthropicPublished 24 July 2026. Listed for completeness and explicitly NOT relied on: the PDF exceeded our fetch size limit, so nothing in it is verified here, including the safety level that secondary outlets attribute to it.
  30. 30Moonshot unveils Kimi K3ForbesIndependent coverage dated 17 July 2026, agreeing with Moonshot's own blog and disagreeing with OpenRouter and Simon Willison, who date the launch 16 July.
  31. 31Alibaba previews Qwen3.8-MaxMarkTechPostIndependent coverage. Source for the withheld activated-parameter count and the absence of a model card, and for the reported multimodal capabilities that no Alibaba page confirms.

Weekly signal · biweekly email

Field-tested ideas. No content treadmill.

One substantial note when we have something worth showing: systems, receipts, prompts, and what failed. Confirm by email. Unsubscribe whenever you like.

Signup opens when the production Turnstile site key is configured.