Search

Tag: #fast-inference6 results

Opus 4.7 Fast Mode on AI Gateway

Opus 4.7 Fast Mode on AI Gateway

Fast mode for Claude Opus 4.7 is now available in research preview on Vercel AI Gateway, delivering ~2.5x faster output token generation with full model intelligence. This experimental feature is priced at 6x standard Opus rates, with input at $30/1M tokens and output at $150/1M tokens. Enable it via provider options or environment variables for Claude Code.

OpenAI Launches Speedy Codex-Spark Model

OpenAI Launches Speedy Codex-Spark Model

OpenAI released GPT-5.3-Codex-Spark, a lightweight version of its Codex intelligent agent programming tool. This slimmed-down model prioritizes extreme inference speed for rapid iteration scenarios. It follows the latest full Codex model released earlier this month.

cnBeta (Full RSS)MediaFeb 13#launch#openai#gpt-53-codex-spark