SourceStalecollected in 7h

Cloudflare Opens Decision Models and RL Fine-Tuning

Read original on Cloudflare Blog
#decision-models#edge-ai

Cloudflare combines open decision models with a practical path to custom reinforcement learning.

30-Second TL;DR

What Changed

Clef and Clef-flash target high-speed classification and agentic workflows.

Why It Matters

The release could make small decision models more practical for latency-sensitive edge and workflow applications. Custom reinforcement learning may help teams adapt models to domain-specific decisions without relying only on prompting.

What To Do Next

Deploy Clef-flash on Workers AI for a representative classification workload and compare latency against your current model.

Who should care:Developers & AI Engineers

Key Points

  • •Clef and Clef-flash target high-speed classification and agentic workflows.
  • •The models are hosted on Cloudflare Workers AI.
  • •A new RL platform supports fine-tuning with developer-provided data.

Deep Insight

Background and context from public sources — not the original article. 7 sources cited.

Enhanced Key Takeaways

  • •Clef consists of a 27B parameter multimodal model and a lightweight 9B parameter Clef-flash variant, both released open-source under an Apache 2.0 license on Hugging Face.
  • •Unlike standard generative LLMs, Clef utilizes a modified Qwen backbone with a non-autoregressive, prefill-only scoring mechanism to yield calibrated probabilities over bounded schemas.
  • •Both models natively support multimodal inputs—including text, JSON, and images—and support up to a 256k token context window (with 64k tokens hosted on Workers AI).
  • •Clef offers drop-in compatibility with Typesafe AI's Jev API schema and outperforms Jev on 3 of 4 Decision Index benchmark categories, with Clef-flash delivering up to ~13× lower latency.
  • •The rollout includes a dedicated RL fine-tuning pipeline initially staffed by forward-deployed engineers, connecting Cloudflare AI Gateway, Containers, and Workers AI.

Competitor Analysis

Cloudflare Clef / Clef-flash
Model Architecture / Modality
27B & 9B parameters; non-autoregressive, prefill-only; Multimodal (text, JSON, vision)
Licensing
Open Source (Apache 2.0)
Context Window
Up to 256k (64k on Workers AI)
Key Performance / Latency
Scores 98.47 on BFCL case-exact; Clef-flash up to ~13× faster on edge GPUs
Typesafe AI Jev (System One)
Model Architecture / Modality
Proprietary decision model; Text-focused
Licensing
Closed Source / Proprietary API
Context Window
32k tokens
Key Performance / Latency
Benchmark baseline (95.75 on BFCL case-exact)
Amazon Strands Decider
Model Architecture / Modality
2B parameter compact decider; Text-focused
Licensing
Open Source
Context Window
Unspecified
Key Performance / Latency
Lightweight footprint targeting edge/micro-classification

Technical Deep Dive

  • Underlying Backbone: Built on a modified Qwen architecture optimized specifically for deterministic classification and decision routing.
  • Inference Mechanism: Uses a non-autoregressive, prefill-only architecture that bypasses sequential token generation, directly scoring and calibrating probability distributions over bounded schema choices.
  • Multimodal Ingestion: Integrates native vision encoders, allowing single-pass classification over text, structured JSON schemas, and image payloads.
  • Context Handling: Architecturally trained to handle context windows up to 256,000 tokens, with hosted edge deployment on Workers AI supporting up to 64,000 tokens.
  • API Interoperability: Implements full API drop-in compatibility with Typesafe AI's Jev schema format, preventing code refactoring during migration.
  • Fine-Tuning Stack: Supported by an integrated RL fine-tuning pipeline orchestrating Cloudflare AI Gateway for trace collection, Cloudflare Containers for training jobs, and Workers AI for edge serving.

Future ImplicationsAI analysis grounded in cited sources

Agentic workflows will shift away from autoregressive generative models for routing and classification tasks.
Prefill-only decision models offer deterministic schemas, calibrated probabilities, and up to 13× lower latency compared to generative token decoding.
Open-source decision models will diminish pricing power for proprietary routing APIs.
Apache 2.0 weights with drop-in compatibility for proprietary APIs like Jev allow organizations to deploy high-performing routing layers locally or on serverless edge infrastructure without vendor lock-in.

Timeline

2026-10
Cloudflare launches Clef and Clef-flash decision models and RL platform during Birthday Week 2026

Weekly AI Recap

Read this week's curated digest of top AI events →

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Cloudflare Blog ↗

This is a summary, not the original. Read the source, or get the weekly briefing.

The weekly digest

One email a week. Unsubscribe anytime.