來源Cloudflare Blog•較早收集於 3h
Cloudflare 內部 AI 堆疊處理 241B 令牌

#ai-engineering#internal-use#edge-inferenceworkers-aicloudflareai-gatewayworkers-ai
💡了解 Cloudflare 如何擴展內部 AI:Workers AI 處理 2410 億令牌。
⚡ 30 秒速覽
有什麼變化
AI Gateway 路由 20 百萬請求
為什麼重要
展現 Cloudflare AI 工具的生產級可靠性,為企業內部 AI 管道提供藍圖。
下一步行動
透過 Cloudflare AI Gateway 路由 AI 請求以實現可擴展推論。
誰應關注:Developers & AI Engineers
關鍵要點
- •AI Gateway 路由 20 百萬請求
- •內部處理 2410 億令牌
- •Workers AI 服務超過 3,683 名內部用戶
🧠 深度解析
本篇為 AI 生成分析,非原文內容。
🔑 增強重點摘要
- •Cloudflare's internal stack utilizes a 'dogfooding' strategy, where the company uses its own production-grade Workers AI and AI Gateway to manage internal LLM workloads, effectively stress-testing its infrastructure at scale.
- •The architecture leverages Cloudflare's global network to perform inference at the edge, reducing latency for internal users by executing models closer to their geographic location rather than relying on centralized data centers.
- •The internal implementation incorporates automated cost-tracking and rate-limiting features via AI Gateway, allowing Cloudflare to monitor and optimize token consumption across various internal departments and AI-driven projects.
📊 競品分析▸ Show
| Feature | Cloudflare (Workers AI/Gateway) | AWS (Bedrock/App Mesh) | Google Cloud (Vertex AI/Gateway) |
|---|---|---|---|
| Primary Focus | Edge-native, low-latency inference | Enterprise-grade, broad model choice | Integrated ML pipeline, deep data stack |
| Pricing Model | Per-token/request, edge-optimized | Per-token/provisioned throughput | Per-token/compute-hour |
| Deployment | Global Edge Network | Regional Data Centers | Regional/Multi-region Data Centers |
🛠️ 技術深入
- •Workers AI utilizes a serverless execution model that dynamically schedules inference tasks across Cloudflare's global fleet of GPUs.
- •AI Gateway acts as a unified proxy layer, providing caching, logging, and analytics for requests sent to both third-party LLM APIs (like OpenAI or Anthropic) and self-hosted models running on Workers AI.
- •The system employs a multi-tenant architecture that isolates internal user workloads while maintaining shared access to model weights cached at the edge.
- •Implementation relies on the Workers runtime, allowing developers to write inference logic in JavaScript/TypeScript that executes directly within the request-response lifecycle.
🔮 前景展望基於引用來源的 AI 分析
Cloudflare will transition to a hybrid-model routing strategy for enterprise customers.
The success of their internal stack demonstrates that routing traffic between edge-hosted models and external APIs is a viable method to balance cost and performance.
Workers AI will expand support for fine-tuned, domain-specific models.
As internal usage grows, the need for specialized models that outperform general-purpose LLMs will drive the development of persistent storage and fine-tuning capabilities within the Workers environment.
⏳ 時間線
2023-09
Cloudflare launches Workers AI to allow developers to run AI models on its global network.
2023-09
Cloudflare introduces AI Gateway to provide observability and control for AI application traffic.
2024-05
Cloudflare expands Workers AI to support Llama 3 and other open-source models at the edge.
2025-02
Cloudflare announces significant performance optimizations for its internal AI inference stack.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Cloudflare Blog ↗
每週電子報
每週一封,可隨時退訂。
