Cloudflare Opens Decision Models and RL Fine-Tuning

Cloudflare combines open decision models with a practical path to custom reinforcement learning.
30-Second TL;DR
What Changed
Clef and Clef-flash target high-speed classification and agentic workflows.
Why It Matters
The release could make small decision models more practical for latency-sensitive edge and workflow applications. Custom reinforcement learning may help teams adapt models to domain-specific decisions without relying only on prompting.
What To Do Next
Deploy Clef-flash on Workers AI for a representative classification workload and compare latency against your current model.
Key Points
- •Clef and Clef-flash target high-speed classification and agentic workflows.
- •The models are hosted on Cloudflare Workers AI.
- •A new RL platform supports fine-tuning with developer-provided data.
Deep Insight
Background and context from public sources — not the original article. 7 sources cited.
Enhanced Key Takeaways
- •Clef consists of a 27B parameter multimodal model and a lightweight 9B parameter Clef-flash variant, both released open-source under an Apache 2.0 license on Hugging Face.
- •Unlike standard generative LLMs, Clef utilizes a modified Qwen backbone with a non-autoregressive, prefill-only scoring mechanism to yield calibrated probabilities over bounded schemas.
- •Both models natively support multimodal inputs—including text, JSON, and images—and support up to a 256k token context window (with 64k tokens hosted on Workers AI).
- •Clef offers drop-in compatibility with Typesafe AI's Jev API schema and outperforms Jev on 3 of 4 Decision Index benchmark categories, with Clef-flash delivering up to ~13× lower latency.
- •The rollout includes a dedicated RL fine-tuning pipeline initially staffed by forward-deployed engineers, connecting Cloudflare AI Gateway, Containers, and Workers AI.
Competitor Analysis
- Model Architecture / Modality
- 27B & 9B parameters; non-autoregressive, prefill-only; Multimodal (text, JSON, vision)
- Licensing
- Open Source (Apache 2.0)
- Context Window
- Up to 256k (64k on Workers AI)
- Key Performance / Latency
- Scores 98.47 on BFCL case-exact; Clef-flash up to ~13× faster on edge GPUs
- Model Architecture / Modality
- Proprietary decision model; Text-focused
- Licensing
- Closed Source / Proprietary API
- Context Window
- 32k tokens
- Key Performance / Latency
- Benchmark baseline (95.75 on BFCL case-exact)
- Model Architecture / Modality
- 2B parameter compact decider; Text-focused
- Licensing
- Open Source
- Context Window
- Unspecified
- Key Performance / Latency
- Lightweight footprint targeting edge/micro-classification
| Competitor / Product | Model Architecture / Modality | Licensing | Context Window | Key Performance / Latency |
|---|---|---|---|---|
| Cloudflare Clef / Clef-flash | 27B & 9B parameters; non-autoregressive, prefill-only; Multimodal (text, JSON, vision) | Open Source (Apache 2.0) | Up to 256k (64k on Workers AI) | Scores 98.47 on BFCL case-exact; Clef-flash up to ~13× faster on edge GPUs |
| Typesafe AI Jev (System One) | Proprietary decision model; Text-focused | Closed Source / Proprietary API | 32k tokens | Benchmark baseline (95.75 on BFCL case-exact) |
| Amazon Strands Decider | 2B parameter compact decider; Text-focused | Open Source | Unspecified | Lightweight footprint targeting edge/micro-classification |
Technical Deep Dive
- Underlying Backbone: Built on a modified Qwen architecture optimized specifically for deterministic classification and decision routing.
- Inference Mechanism: Uses a non-autoregressive, prefill-only architecture that bypasses sequential token generation, directly scoring and calibrating probability distributions over bounded schema choices.
- Multimodal Ingestion: Integrates native vision encoders, allowing single-pass classification over text, structured JSON schemas, and image payloads.
- Context Handling: Architecturally trained to handle context windows up to 256,000 tokens, with hosted edge deployment on Workers AI supporting up to 64,000 tokens.
- API Interoperability: Implements full API drop-in compatibility with Typesafe AI's Jev schema format, preventing code refactoring during migration.
- Fine-Tuning Stack: Supported by an integrated RL fine-tuning pipeline orchestrating Cloudflare AI Gateway for trace collection, Cloudflare Containers for training jobs, and Workers AI for edge serving.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2026-10Cloudflare launches Clef and Clef-flash decision models and RL platform during Birthday Week 2026
Sources (7)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Cloudflare Blog ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.