
COMPASS: Grounding Composition-Intent in Unified Multimodal Models
COMPASS is a new unified multimodal framework that improves fine-grained composition recognition and control. By using a shared expert token for both perception and generation, it enables precise layout control in image synthesis.
ArXiv AI · 81d ago

















