Together AI adds Thinking Machines Lab’s Inkling model

Access the latest Thinking Machines Lab model instantly via Together AI's high-performance inference API.
30-Second TL;DR
What Changed
Inkling model is now available on Together AI platform
Why It Matters
This integration enables developers to rapidly prototype and deploy the latest research models without managing infrastructure. It lowers the barrier to entry for testing new, cutting-edge architectures.
What To Do Next
Visit the Together AI model catalog to test Inkling via the API and compare its performance against existing benchmarks.
Key Points
- •Inkling model is now available on Together AI platform
- •Day 0 availability for the latest model from Thinking Machines Lab
- •Expands the library of models accessible via Together AI's inference API
Deep Insight
AI-generated analysis for this event — not the original article.
Enhanced Key Takeaways
- •Inkling is specifically optimized for low-latency reasoning tasks, distinguishing it from general-purpose large language models.
- •The model utilizes a novel 'Chain-of-Thought Distillation' architecture developed by Thinking Machines Lab to reduce inference costs.
- •Together AI's integration includes support for fine-tuning Inkling on proprietary datasets via their serverless fine-tuning API.
- •Thinking Machines Lab positions Inkling as a specialized alternative to larger models for edge-computing and real-time agentic workflows.
- •The partnership marks the first time Thinking Machines Lab has utilized a third-party inference provider for their flagship model release.
Competitor Analysis
- Inkling (Together AI)
- Low-latency Reasoning
- Groq (Llama 3.1)
- Raw Inference Speed
- Fireworks AI (Qwen 2.5)
- High-throughput Serving
- Inkling (Together AI)
- Competitive per-token
- Groq (Llama 3.1)
- Aggressive/Volume-based
- Fireworks AI (Qwen 2.5)
- Tiered/Enterprise
- Inkling (Together AI)
- CoT Distilled
- Groq (Llama 3.1)
- Standard Transformer
- Fireworks AI (Qwen 2.5)
- Standard Transformer
| Feature | Inkling (Together AI) | Groq (Llama 3.1) | Fireworks AI (Qwen 2.5) |
|---|---|---|---|
| Primary Focus | Low-latency Reasoning | Raw Inference Speed | High-throughput Serving |
| Pricing | Competitive per-token | Aggressive/Volume-based | Tiered/Enterprise |
| Architecture | CoT Distilled | Standard Transformer | Standard Transformer |
Technical Deep Dive
- Model Architecture: Employs a sparse mixture-of-experts (MoE) backbone with a dedicated reasoning head for intermediate step generation.
- Context Window: Supports a 128k token context window with optimized KV-caching for long-sequence reasoning.
- Quantization: Native support for FP8 and INT4 inference modes to maximize throughput on H100/A100 clusters.
- Training Methodology: Utilized synthetic data generation techniques to refine reasoning traces during the post-training phase.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2025-03Thinking Machines Lab founded with a focus on reasoning-centric AI architectures.
- 2025-11Thinking Machines Lab releases the first research paper on Chain-of-Thought Distillation.
- 2026-05Together AI announces expanded support for specialized reasoning models on their platform.
- 2026-07Inkling model officially launched and integrated into Together AI's inference API.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Together AI Blog ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.