FR
live
AI

Cloudflare and AWS open the open-weight decision-model lane for agents

On 1 October 2026 Cloudflare released Clef and Clef-flash, and AWS Strands Decider 2B: open-source decision models that answer with typed probabilities instead of text, built for agent decision-making. They undercut the closed Jev model and raise the question of what you should stop sending to the frontier model.

A single piano key glowing warm amber among rows of dark, unlit piano keys in a dimly lit concert hall.

1 October 2026. Cloudflare releases Clef (27B) and Clef-flash (9B), its first in-house-trained models, under an Apache 2.0 license. 1 October 2026. AWS Strands Labs opens the weights of Strands Decider 2B, a 2-billion-parameter decision model built to run locally. 2 October 2026. Digital Applied’s tally confirms that no frontier model opened the month: specialist models — decision, voice, translation — are the headline. Why it matters: a new model category, the decision model, is moving out of the closed format and becoming a standard open-source component in the agent loop.

What a decision model is, and what it is not

A decision model does not generate prose. Where an LLM produces open-ended, non-deterministic text, a decision model receives a state (a ticket, a domain, a workflow step) and returns typed answers with probabilities: a choice from a list, a score, a boolean. The agent’s code reads that structured answer to route a ticket, trigger an escalation, or defer to a human.

The difference is operational, not cosmetic. Cloudflare gives the example of its Threat Intelligence team: to classify a domain, Clef took 2.2 seconds to load, render and classify the page, where its fastest general LLM, gpt-oss-120b, took 4.7 and returned only two classifications. Half the latency, and outputs the code can consume directly — no recall prompt, no fragile parser, no retry on an out-of-schema answer.

The call is trivial and fully compatible with the API of Jev, the closed TypeSafe model that launched the category:

json
{
  "model": "clef",
  "state": "Checkout has been failing for every customer for the last hour.",
  "questions": {
    "team": {
      "type": "choice",
      "instructions": "Which team should handle this request?",
      "criteria": {
        "billing": "Payments, invoices, and refunds",
        "technical": "Outages, errors, and configuration"
      }
    }
  }
}

Clef versus Jev: faster, multimodal, more context

Clef stands apart from Jev on three concrete points. First, a vision encoder: it accepts images and video, where Jev classifies text only. Second, a 64,000-token context window, against 32,000 for Jev — room to inject more state to classify. Third, latency: Clef posts a median of 209.3 ms and Clef-flash of 38.8 ms, against 524.1 ms for Jev. Cloudflare’s changelog sums up the gap as “up to 13x faster than Jev”.

On quality, Clef leads the Jev Decision Index: 98.47% case-exact on BFCL, 94.20% macro-F1 on BANKING77, 97.43% on CLINC150+OOS. Clef-flash wins several tests despite its small size — 98.76% on BFCL, 93.11% on API-Bank. Both are hosted on Workers AI at $0.24 and $0.09 per million input tokens, and Cloudflare commits to not reading, storing, or training on your requests unless you opt into its upcoming reinforcement-learning fine-tuning.

Decider 2B, AWS’s local counterpart

AWS’s answer points the other way: Strands Decider 2B is built to run locally, on CPU or GPU, with the weights, training data, and scripts all published. Its contract is minimal: it scores predefined options and returns a choice or a confidence value in a single forward pass — about 115 ms on an RTX 3090, according to the reported measurements. Where Clef targets decision-making in the server-side critical path, Decider 2B targets local development, experimentation, and zero marginal cost.

Both releases tell the same story: the “decide fast, cheaply, with a guaranteed output” lane is filling up, and it is filling up in open weights. Of the eight models released between 1 and 3 October, four shipped under Apache 2.0 with public weights — Clef, Clef-flash, Decider 2B, and Bilibili’s Index-Translate-35B-A3B translation model.

The 1 October field: specialization becomes the norm

Clef and Decider 2B are not alone. On 1 October, Microsoft AI released three speech models — MAI-Transcribe-2-Streaming, continuous transcription covering 60 languages, and the two MAI-Voice-2.1 text-to-speech models in 23 languages with a Flash variant claimed at 150 ms end to end — while Tavus opened early access to Griffin-Lite, a full-duplex video conversation model. On 2 October, Bilibili’s Index team published Index-Translate-35B-A3B, a mixture-of-experts translator with 35 billion total and 3 billion active parameters, spanning 150 languages.

The regularity is the message. After a September that opened with three frontier models in two days, October begins with models that do one thing: decide, transcribe, synthesize, translate, converse. And two of the five vendors, Cloudflare and Amazon, publish weights rather than selling tokens. Of the eight models released between 1 and 3 October, four shipped under Apache 2.0 with public weights — half of the opening month is open source, a shift that was not a given a year ago.

That specialization redraws the economics. The reflex of “everything goes through the frontier model” costs money and adds latency for tasks that need no reasoning. A decision model at $0.09 per million tokens, or a local Decider 2B at zero marginal cost, shifts the load to the cheapest layer — provided the typed output is reliable, which always circles back to calibration.

Where it sits in the agent

The buying question is not “is it better than the frontier model” but “what does it let me stop sending to the frontier model”. A reasoning LLM is expensive and non-deterministic; having it decide every branch of a workflow is waste. A decision model slots in just before the tool call: it routes, classifies, triggers, or withholds, for a fraction of the cost and with predictable latency.

The risk not to underestimate is calibration: a displayed probability is only trustworthy if the model is calibrated on your real distributions, not on generic benchmarks. Clef posts strong scores on public sets, but an agent’s decision over production data has to be validated on your own inputs — which is exactly what Cloudflare is pushing with its RL fine-tuning product, and what the open format finally makes verifiable locally.

There is also a strategic signal in who is shipping these models. Cloudflare and Amazon both run inference infrastructure, yet both chose to publish weights under Apache 2.0 rather than gate the models behind an API. That runs opposite to the lock-in playbook: a decision model is cheap enough to run anywhere, so the value migrates to the surrounding tooling — hosting, fine-tuning, observability — rather than to the weights themselves. For buyers, it means the category is unlikely to become the kind of closed garden that frontier models became.

Verdict

If you build agents that decide in the critical path, try Clef-flash for routing and classification: at $0.09 per million tokens and 38.8 ms median latency, it moves decisions off the frontier model without breaking your budget. If you develop locally or want to audit the weights, Strands Decider 2B is the best entry point: everything is published, and it runs on a plain CPU. If you are already on Jev, the API compatibility makes the swap immediate, and open source removes the dependency risk on a closed model. In every case, validate calibration on your own data before letting a decision model take autonomous actions: a probability is only a guarantee if it was measured on your distribution, not on a leaderboard.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

Reflection AI opens Beam, a 501B model that challenges Chinese open weights

On October 5, 2026, Reflection AI unveiled Beam, a 501-billion-parameter open-source model trained in eight weeks on NVIDIA hardware rented from SpaceX, and which approaches the performance of two-trillion-parameter models. For teams evaluating open-weight models, Beam shifts the cost-performance calculus, but its weights only land at the end of the month and still trail closed frontier models.

Google readies Gemini to control the entire Mac through a hidden sandbox option

On 3 October 2026, BleepingComputer reported that Google is testing desktop control for Gemini, with a hidden “Additional sandbox options” setting that would let the AI read, create, modify or delete files and drive Mail, Safari and Messages. The feature is not live yet, and Apple is weighing whether to make it harder for AI agents to reach personal files on the Mac.

← Back to the feed

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss