
TTE-Flash: Efficient Reasoning-Aware Multimodal Representations
TTE-Flash introduces 'latent think tokens' to replace explicit Chain-of-Thought generation, significantly reducing computational overhead for multimodal embeddings. The model achieves superior performance on the MMEB-v2 benchmark while maintaining constant inference costs.


