來源虎嗅•較早收集於 50m
銀行業面臨Token高昂成本的ROI壓力
💡了解大型銀行如何遏制失控的LLM成本,並轉向以ROI為中心的AI部署。
⚡ 30 秒速覽
有什麼變化
大型銀行日均Token消耗量達數十億級別,顯著增加了IT成本。
為什麼重要
這標誌著企業AI採用進入冷靜期,單純的Token使用量等「虛榮指標」將被具體的商業價值指標所取代。
下一步行動
為您的LLM應用開發細粒度的ROI追蹤儀表板,以便向利益相關者證明基礎設施成本的合理性。
誰應關注:Founders & Product Leaders
關鍵要點
- •大型銀行日均Token消耗量達數十億級別,顯著增加了IT成本。
- •銀行業重心正從「AI採用」轉向「AI效率」與ROI衡量。
- •財富管理與風控領域中效果不佳的AI智能體正面臨預算削減。
🧠 深度解析
本篇為 AI 生成分析,非原文內容。
🔑 增強重點摘要
- •Chinese financial institutions are increasingly pivoting toward 'Small Language Models' (SLMs) and domain-specific fine-tuning to reduce dependency on high-cost, general-purpose frontier models.
- •The 'token inflation' crisis is driving a shift toward hybrid AI architectures where simple, rule-based automation handles routine queries, reserving expensive LLM calls for complex reasoning tasks.
- •Regulators in China have begun emphasizing 'AI cost-transparency' in financial audits, requiring banks to justify the energy and compute expenditure of AI deployments against tangible productivity gains.
- •Major banks are renegotiating cloud-compute contracts, moving away from pay-per-token models toward reserved capacity or private-cloud deployments to stabilize unpredictable IT expenditures.
- •Internal data shows that 'agentic workflows'—where AI agents autonomously chain multiple calls—are the primary drivers of the observed token explosion, prompting a move toward human-in-the-loop verification for high-cost processes.
📊 競品分析▸ Show
| Feature | General-Purpose LLMs (e.g., GPT-4/Claude) | Domain-Specific SLMs (e.g., Qwen-Finance/DeepSeek) | Rule-Based Automation |
|---|---|---|---|
| Token Cost | Extremely High | Low to Moderate | Negligible |
| Reasoning Capability | Superior | Moderate | None |
| Deployment | Public API / Cloud | Private / On-Premise | On-Premise |
| Latency | High | Low | Instant |
🛠️ 技術深入
- Shift toward Mixture-of-Experts (MoE) architectures to activate only necessary parameters per query, reducing total compute per token.
- Implementation of 'Prompt Caching' techniques to store frequently used context, significantly lowering input token costs for recurring wealth management queries.
- Adoption of Knowledge Graph-augmented generation (GraphRAG) to improve accuracy in risk control, reducing the need for multiple 'retry' tokens caused by model hallucinations.
- Transition to quantized models (INT8/INT4) for internal deployment to maximize throughput on existing GPU clusters without sacrificing critical financial reasoning accuracy.
🔮 前景展望基於引用來源的 AI 分析
Banks will mandate 'Token-per-Task' budgets for all AI development teams by Q4 2026.
The current uncontrolled consumption model is unsustainable, forcing financial institutions to treat AI compute as a strictly rationed resource similar to headcount.
The market share of proprietary, in-house trained models will surpass general-purpose API usage in Chinese banking by 2027.
The need for cost control and data sovereignty is making the long-term ROI of training smaller, specialized models more attractive than perpetual API fees.
⏳ 時間線
2023-05
Initial wave of generative AI pilot programs launched across major Chinese state-owned banks.
2024-02
Rapid scaling of AI agents in customer service and wealth management leads to first reports of unexpected IT budget overruns.
2025-09
Financial regulators issue preliminary guidance on AI operational risk, highlighting the need for cost-efficiency and model stability.
2026-03
Major banks initiate 'AI Efficiency' audits, freezing funding for high-token-consumption projects with unproven ROI.
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 虎嗅 ↗
每週電子報
每週一封,可隨時退訂。



