Harmonizer enabled for top-policy reruns
The Harmonizer diffusion model will be disabled for initial runs. It will be enabled when organizers perform the final runs of the top policies.
Welcome to the AlpaSim End to End Closed Loop Challenge. This competition invites teams to build autonomous driving policies and compare them head-to-head in realistic closed-loop simulation, where each policy's decisions shape the future scene it must handle.
Autonomous driving research has made major progress, but it remains hard to compare policies across labs and companies in a realistic, reproducible way. Open-loop evaluation is useful, but it misses the compounding effects that make driving hard: a small planning error can change future observations, interactions, and risk.
AlpaSim provides a shared simulator, public development data, starter tools, baseline policies, and a common containerized submission interface. Organizer-managed evaluation workers run submissions on private held-out scenarios and publish leaderboard results with both a policy capability score and a safety metric.
Across both tracks, the goal is not just to crown a winner. We want to learn which policy families are robust under distribution shift, where they fail, and how the community can make AV evaluation more trustworthy.
The Harmonizer diffusion model will be disabled for initial runs. It will be enabled when organizers perform the final runs of the top policies.
cd713e0d0563352ce45b99d0a53a7173d581bde2
This is the current competition drop on the e2e_challenge
branch. Its active configuration includes the settings below.
The submission limit is now three per team.
PAI evaluations use the final, larger evaluation suite and may take nearly an order of magnitude longer to complete.
Teams can now provide values for tunable controller parameters at submission time.
Per internal legal guidance, terms and conditions have been added for competitors. All registered teams may submit, but some teams may be ineligible for prizes. After registering and before submitting, please read and accept the terms through the challenge command-line tool.
PAI rollouts now include rear-facing cameras, giving drivers image coverage around the full ego vehicle rather than only its forward and side views.
The rollout specification now includes the ego vehicle bounding box and its pose relative to the ego rig, so drivers can account for the vehicle footprint used by evaluation.
New local evaluation utilities include a curated PAI validation split and reference runs, so teams can construct a local ranking and better understand the evaluation criteria. The local score scale is illustrative and is not directly comparable to the official leaderboard.
Teams can now run the full public navtest suite
locally for the nuPlan track. It is a public development suite,
not the private official competition suite, and no nuPlan
reference bundle is included yet.
No known issues remain in the current drop. Let us know on the AlpaSim GitHub issues page if you find anything.
6b24e0159e5f92a1d9150ddab81f933745e24758
This previous version is retained for reference. It was also
on the e2e_challenge branch.
Some soft-open scenes could produce violations even when following
the ground-truth path. This was fixed in the current server drop,
cd713e0d0563352ce45b99d0a53a7173d581bde2.
Registration is limited to teams affiliated with academic or industry organizations. To protect limited evaluation capacity and ensure fair participation, the team captain's verified Hugging Face primary email must be an organization address.
Registrations are reviewed to prevent duplicate or non-genuine accounts from circumventing team-level participation limits. Independent groups from the same organization are welcome: identify your lab, department, or group and briefly describe the project in the registration form.
Full challenge rules and submission limits are in the Contestant Guide.
The challenge is designed to compare independently developed policies under a shared, organizer-managed evaluation protocol. These rules protect that comparison and the limited evaluation capacity available to all teams.
Each participant may belong to only one registered team. Submission limits apply at the team level across its authorized submitters; teams and participants must not coordinate submissions or share accounts to bypass those limits. Multiple independent teams from the same organization are welcome.
Award-eligible nuPlan submissions may use only publicly available datasets, annotations, and model weights. PAI-AV submissions may use private resources when the team has the right to use them and provides a high-level disclosure of their use.
Submissions must not use hidden scenario information, private route identifiers, or any information unavailable to the policy during evaluation. Attempts to exploit the leaderboard, simulator, or platform; attack the service; or otherwise undermine fair evaluation may result in disqualification.
Submitted policies must follow the published container contract and evaluation constraints. Organizers may request code, logs, or method details for award candidates, disputed results, or suspected rule violations, and may invalidate non-reproducible submissions.
Award eligibility requires a final valid submission and a concise technical report for the relevant track. Reports are limited to four pages, excluding references and optional appendices, and should use the AlpaSim technical report template, which uses an attributed adaptation of the NeurIPS 2026 layout. Innovation prizes will be determined through organizer review of these reports and may overlap with leaderboard placement awards.
Use the AlpaSim GitHub issues for questions, reproducible bugs, and rule clarifications. Material clarifications will be reflected here.
The competition has two complementary tracks, covering both large-scale geographically diverse driving data and a lower-barrier entry point for teams working with established nuPlan-style workflows.
The larger-scale setting for testing whether promising policies hold up as scenario diversity and long-tail coverage increase.
Competition scenes come from Denmark, France, Germany, Japan, South Korea, Spain, Sweden, the United Kingdom, and the United States.
A lower-barrier track for teams building on the widely used nuPlan ecosystem or NAVSIM-style development workflows.
The nuPlan track uses a private evaluation dataset.
Each track will award two NVIDIA DGX Spark prizes: one for first place and one for an innovative solution.
The winners will receive a DGX Spark with an approximate retail value of USD $4,700. The prize must be accepted as awarded and is non-transferable, non-exchangeable, and has no cash value. NVIDIA is not responsible for any costs or taxes associated with prize acceptance.
Two NVIDIA DGX Spark prizes awarded.
Two NVIDIA DGX Spark prizes awarded.
Challenge setup, tracks, scoring, and participant workflow.
Baseline files and examples for building an initial submission.
Command-line tooling for packaging, validating, and submitting entries.
Container image constraints and requirements for valid submissions.
Reference documentation for how policy capability scores are computed.
Open-source training and closed-loop evaluation framework with pretrained baselines for the AlpaSim E2E Challenge.
Four-page report skeleton for award-eligible final submissions, using an attributed NeurIPS 2026-style layout.
Use the public AlpaSim repository for bug reports, questions, and competition discussions.
This competition is hosted by NVIDIA's Autonomous Vehicle Research Group, KE:SAI, and HKU.
Interdisciplinary NVIDIA Research team advancing vehicle autonomy across perception, prediction, planning, control, simulation, foundation models, and AI safety.
Non-profit open-science research lab advancing robust, safe, and reproducible physical AI, with a focus on world models, autonomy, and open self-driving technology.
OpenDriveLab at The University of Hong Kong is a research group focused on embodied AI and autonomous driving. The lab develops open benchmarks, simulation infrastructure, world models, and end-to-end driving systems to support scalable, reproducible, and trustworthy physical AI research.