Superhuman Generals.io agent built with self-play RL
Learn how to scale RL agents in RTS games using JAX and Vision Transformers for superhuman performance.
30-Second TL;DR
What Changed
Achieved #1 ranking on the human 1v1 leaderboard using self-play RL.
Why It Matters
Demonstrates the effectiveness of scaling-first approaches in complex, imperfect-information RTS environments. Provides a valuable open-source framework for researchers working on game-based AI.
What To Do Next
Clone the repository and experiment with the JAX-based simulator to test your own RL agents in an imperfect-information RTS environment.
Key Points
- •Achieved #1 ranking on the human 1v1 leaderboard using self-play RL.
- •Reimplemented the training pipeline in JAX for significant performance gains.
- •Switched from CNN to Vision Transformer to prioritize scaling over human priors.
- •Open-sourced the JAX-based simulator and agent code for community use.
Deep Insight
AI-generated analysis for this event — not the original article.
Enhanced Key Takeaways
- •The agent utilizes a custom-built, vectorized environment in JAX that allows for thousands of parallel game simulations, significantly accelerating the training throughput compared to standard Python-based environments.
- •The Vision Transformer (ViT) architecture was specifically chosen to handle the game's grid-based state representation as a sequence of patches, enabling the model to learn spatial relationships without the inductive biases inherent in CNNs.
- •The project addresses the 'sparse reward' problem in Generals.io by implementing a multi-stage reward shaping strategy that incentivizes early-game expansion and mid-game unit efficiency.
- •Training was conducted using a distributed PPO (Proximal Policy Optimization) implementation, which proved critical for stabilizing the policy updates during the intense self-play phase.
- •The agent's superhuman performance is attributed to its ability to discover 'non-human' strategies, such as hyper-aggressive fog-of-war exploitation that human players struggle to counter.
Technical Deep Dive
- Architecture: Vision Transformer (ViT) backbone with a custom patch embedding layer designed for 2D grid inputs.
- Simulation Engine: Custom JAX-based environment providing hardware-accelerated state transitions and observation generation.
- Training Algorithm: Distributed Proximal Policy Optimization (PPO) with generalized advantage estimation (GAE).
- Hardware Utilization: Optimized for TPU/GPU clusters, achieving high throughput by minimizing CPU-GPU data transfer bottlenecks.
- State Representation: Multi-channel tensor input representing unit counts, terrain types, and fog-of-war status.
Future ImplicationsAI analysis grounded in cited sources
Timeline
- 2025-09Initial development of the JAX-based Generals.io simulation environment begins.
- 2026-02Transition from CNN-based architecture to Vision Transformer for policy network.
- 2026-05Agent achieves #1 ranking on the human 1v1 leaderboard.
- 2026-06Project code and simulator open-sourced to the community.
Weekly AI Recap
Read this week's curated digest of top AI events →
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/MachineLearning ↗
This is a summary, not the original. Read the source, or get the weekly briefing.
The weekly digest
One email a week. Unsubscribe anytime.