Understand, predict, simulate, create, plan, and act in the real world.
DreamX is the unified spatial intelligence model and system portfolio of AMAP-ML, the AI team at Alibaba AMAP. We connect research, engineering, and real-world deployment across maps, mobility, local services, digital content, and interactive worlds.
We advance DreamX through production systems, open-source projects, benchmarks, and publications at ICLR, CVPR, ECCV, ACL, AAAI, SIGGRAPH, ICCV, ICML, KDD, EMNLP, ACM MM, and WWW. We release code and evaluation assets to help the community reproduce, compare, and extend our work.
Join us | Research interns, full-time researchers, and AI engineers in spatial intelligence, LLM agents, reinforcement learning, world models, multimodal learning, embodied AI, recommendation, and generative AI are welcome to get in touch.
30+ Open-source Projects |
11 ICLR 2026 Papers |
10 CVPR 2026 Papers |
5 ECCV 2026 Papers
7 ACL 2026 Papers |
5 AAAI 2026 Papers |
4 ICML 2026 Papers |
1 KDD 2026 Oral Paper
5 ICCV 2025 Papers |
2 EMNLP 2025 Oral Papers
Focus: Spatial Intelligence · DreamX · LLM Agents · World Models · Multimodal and Generative AI
We define spatial intelligence as the ability of AI to understand the real world and its evolution over space and time; predict future states; generate and simulate digital representations; plan toward human goals; and act through products and embodied systems.
For AMAP, this means connecting maps, mobility, urban environments, local services, digital content, and physical action in one learning and deployment loop. Generative intelligence is an essential part of this mission: it enables AI to represent, create, enrich, and simulate the world rather than standing as a separate product anchor.
Our work is organized around three core problems:
| Core problem | Primary DreamX families | What it means |
|---|---|---|
| Understand and Predict the World | DreamX-Predictor · DreamX-REC | Connect maps, vision, language, mobility, urban scenes, user intent, and product signals to understand spatial context and forecast how the world evolves. |
| Generate and Simulate the World | DreamX-World · DreamX-Creator | Create map-native assets, videos, 3D scenes, digital content, and persistent interactive worlds with controllability, spatial consistency, temporal coherence, and physical plausibility. |
| Plan and Act in the World | DreamX-Agent · DreamX-Phi | Build agents and decision systems that reason, use tools, plan, self-reflect, and turn human goals into actions in digital and physical environments. |
DreamX turns this mission into six model and system families. Each family focuses on a distinct relationship between intelligence and the real world:
| Family | Role | Focus |
|---|---|---|
| DreamX-Predictor | Predict the world | Model the spatiotemporal evolution of traffic, mobility, demand, supply, and urban conditions. |
| DreamX-World | Simulate the world | Learn dynamic world models for persistent, controllable, physically grounded, and interactive simulation. |
| DreamX-Agent | Plan and complete digital tasks | Understand user goals, reason over spatial context, use tools, and coordinate complex map, mobility, and local-service workflows. |
| DreamX-Phi | Act in the physical world | Connect perception, reasoning, decision-making, and physical action for embodied intelligence, autonomous systems, and spatial devices. |
| DreamX-REC | Connect people, places, and services | Match intent with locations, content, routes, and services under spatial, temporal, and contextual constraints. |
| DreamX-Creator | Create and enrich digital assets | Generate and edit map and navigation assets, local-service content, images, videos, 3D assets, and other spatial media. |
This table defines the portfolio at the family level. Public same-name releases and related research artifacts are linked below where available.
The six DreamX families build on a common foundation:
| Shared capability | Role |
|---|---|
| Spatial data and knowledge | Ground models in maps, routes, places, mobility, urban environments, local services, and real-world feedback. |
| Multimodal foundation models | Connect language, vision, video, maps, GUIs, sensor observations, user intent, and product signals. |
| Spatiotemporal and world modeling | Represent geometry, dynamics, long-horizon evolution, causality, interaction, and physical constraints. |
| Agents, reinforcement learning, and decision-making | Train models to reason, use tools, recommend, plan, self-reflect, and improve through feedback. |
| Generative modeling | Create and edit controllable, consistent, high-quality spatial assets, media, scenes, and experiences. |
| Infrastructure and evaluation | Support scalable data, training, inference, deployment, benchmarks, metrics, and reproducible evaluation. |
These selected public projects show how the spatial intelligence mission translates into research artifacts, open-source systems, benchmarks, and models. The complete project map below organizes all public artifacts by the three core problems and the shared foundation.
| Project | Contribution | Why it matters |
|---|---|---|
| SkillClaw | Agentic skill evolution from real interaction traces. | Demonstrates reusable, self-evolving agent capabilities. |
| FluxText | Scene-text editing for controllable visual asset generation. | Connects generative modeling with practical content creation. |
| Code2World | GUI world modeling through renderable code generation. | Explores executable and interactive world representations. |
| Tree-GRPO | Tree-search rollouts for LLM agent reinforcement learning. | Advances exploration and reasoning in agent training. |
| GPG | Minimalist group policy gradient for model reasoning. | Provides a simple and reusable reinforcement-learning foundation. |
| MobilityBench | Route-planning agent evaluation in real-world mobility scenarios. | Grounds spatial reasoning in an AMAP-native benchmark. |
- 2026.06.18 AMAP-ML has 5 papers accepted to ECCV 2026, expanding the team's work across spatial intelligence, generative modeling, and multimodal AI.
- 2026.06.15 DreamX-World releases its 1.0 technical report and open-sources DreamX-World-5B for long-horizon interactive world generation with 1-minute video support.
- 2026.05.18 MobilityBench provides a scalable benchmark for evaluating route-planning agents in real-world mobility scenarios (KDD 2026 Oral).
- 2026.05.12 CoEvolve trains LLM agents through agent-data mutual evolution, using failure signals to synthesize harder tasks as the agent improves (ACL 2026).
- 2026.05.12 Thinking-with-Map strengthens geolocalization with a reinforced parallel map-augmented reasoning agent (ACL 2026 Findings).
- 2026.05.11 DreamX-World releases the 5B-Cam model and inference code for general-purpose interactive world simulation.
- 2026.05.01 UniMRG shows that multi-representation generation strengthens understanding in unified multimodal models, not just generation (ICML 2026).
- 2026.05.01 Train-Free Infinite-Frame Generation extends pretrained video diffusion to arbitrarily long, temporally consistent videos without any retraining (ICML 2026).
- 2026.05.01 D-Evo improves data efficiency in RL with dual difficulty-aware self-evolution that adaptively reshapes both task and sample difficulty (ICML 2026).
- 2026.05.01 EEPO introduces embedding-perturbed exploration for preference optimization in flow models, addressing exploration collapse in continuous generative policies (ICML 2026).
- 2026.04.22 DCW mitigates SNR-t bias and improves diffusion generation quality across model families (CVPR 2026).
- 2026.04.22 EMF extends efficient one-step generation from class-conditioned synthesis to text-conditioned image generation (CVPR 2026).
- 2026.04.10 SkillClaw turns real interaction traces into reusable, evolving skill libraries.
- 2026.04.01 MACE-Dance decouples motion generation and appearance synthesis for high-quality music-driven dance video (SIGGRAPH 2026).
- 2026.03.23 Omni-WorldBench evaluates world models in dynamic 4D interactive settings.
- 2026.03.20 AutoDrive-R2 improves VLA models with reasoning and self-reflection for autonomous driving scenarios (ICLR 2026).
- 2026.03.18 Video-STAR uses tool-augmented reinforcement learning for open-vocabulary action recognition in video (ICLR 2026).
- 2026.03.11 RL3DEdit uses geometry-guided reinforcement learning to make 3D scene edits more multi-view consistent (CVPR 2026).
Earlier Updates
- 2026.03.01 FE2E transfers image-editing priors into dense depth and normal estimation (CVPR 2026).
- 2026.02.28 FASA improves sparse decoding with frequency-aware attention (ICLR 2026).
- 2026.02.27 Eevee provides high-resolution data and evaluation for video-based virtual try-on (CVPR 2026 Findings).
- 2026.02.06 MobilityBench evaluates route-planning agents in real-world mobility scenarios (KDD 2026 Oral).
- 2026.02.06 SpatialGenEval benchmarks spatial intelligence in text-to-image models (ICLR 2026).
- 2026.02.06 Tree-GRPO replaces independent chain rollouts with tree-search rollouts for LLM agent reinforcement learning (ICLR 2026).
- 2026.02.04 Code2World predicts GUI transitions through renderable code generation.
- 2026.02.04 GPG provides a simple group policy gradient baseline for model reasoning (ICLR 2026).
- 2025.10.22 Taming-Hallucinations reduces MLLM video hallucinations with counterfactual video generation.
- 2025.06.20 FluxText provides a diffusion transformer baseline for scene-text editing.
The project map complements the family-level view with the research and open-source foundations behind DreamX. Each artifact is organized by the primary role it plays in the spatial intelligence stack.
| Repository | Contribution | Venue |
|---|---|---|
| MobilityBench | Route-planning agent evaluation in real-world mobility scenarios. | KDD 2026 Oral |
| Thinking-with-Map | Map-augmented geolocalization agent trained with reinforcement learning. | ACL 2026 Findings |
| SocioReasoner | Vision-language reasoning for urban socio-semantic segmentation. | ICLR 2026 |
| DSFNet | Multi-scenario route ranking with a public industrial driving-route dataset and AMAP deployment. | WWW 2025 |
| IntTravel | Real-world dataset and generative framework for integrated multi-task travel recommendation. | arXiv 2026 |
| FE2E | Image-editing priors for dense geometry estimation. | CVPR 2026 |
| UniVG-R1 | Reasoning-guided universal visual grounding with reinforcement learning. | CVPR 2026 |
| Taming-Hallucinations | Counterfactual video generation for reducing MLLM video hallucinations. | - |
| Repository | Contribution | Venue |
|---|---|---|
| DreamX-World | General-purpose world model for interactive world simulation. | - |
| Code2World | GUI world model via renderable code generation. | - |
| FluxText | Diffusion transformer baseline for scene-text editing. | - |
| RL3DEdit | Geometry-guided reinforcement learning for multi-view consistent 3D scene editing. | CVPR 2026 |
| MACE-Dance | Motion-appearance cascaded generation for music-driven dance video. | SIGGRAPH 2026 |
| Omni-Effects | Prompt-guided and spatially controllable composite visual effects generation. | AAAI 2026 |
| S2-Guidance | Training-free stochastic self-guidance for diffusion models. | ICLR 2026 |
| EPG | Pixel-space generative modeling via self-supervised pre-training. | ICLR 2026 |
| USP | Unified self-supervised pretraining in VAE space for diffusion models. | ICCV 2025 |
| EMF | Text-conditioned one-step image generation. | CVPR 2026 |
| DCW | Differential correction for SNR-t bias in diffusion probabilistic models. | CVPR 2026 |
| NarrLV | Narrative-centric evaluation for long video generation models. | ICLR 2026 |
| ImagerySearch | Adaptive test-time search for video generation. | AAAI 2026 |
| Eevee | High-resolution benchmark for video-based virtual try-on. | CVPR 2026 Findings |
| VMBench | Perception-aligned benchmark for video motion generation. | ICCV 2025 |
| Repository | Contribution | Venue |
|---|---|---|
| SkillClaw | Agentic evolver for collective skill library improvement. | - |
| AutoDrive-R2 | Reasoning and self-reflection for VLA models in autonomous driving. | ICLR 2026 |
| Tree-GRPO | Tree-search rollouts for LLM agent reinforcement learning. | ICLR 2026 |
| GPG | Simple and strong group policy gradient baseline for model reasoning. | ICLR 2026 |
| CoEvolve | Agent-data mutual evolution for training LLM agents. | ACL 2026 |
| MathForge | Difficulty-aware GRPO and multi-aspect reformulation for math reasoning. | ICLR 2026 |
| Video-STAR | Tool-using reinforcement learning for open-vocabulary action recognition. | ICLR 2026 |
| Repository | Contribution | Venue |
|---|---|---|
| SpatialGenEval | Spatial intelligence evaluation for text-to-image models. | ICLR 2026 |
| Omni-WorldBench | Benchmark for interactive response capabilities of world models. | arXiv 2026 |
| RealQA | Realistic image quality and aesthetic scoring with multimodal LLMs. | - |
| FASA | Frequency-aware sparse attention for efficient sparse decoding. | ICLR 2026 |
We are looking for people who want to build the next generation of spatial intelligence systems through clean code, reproducible experiments, rigorous evaluation, ambitious problem selection, and real-world product impact.
If you are interested in research internships, full-time roles, or academic collaboration, please email cxxgtxy@gmail.com (homepage) with your CV, representative projects, and research interests.