long-horizon agent tasks · signed
Agent leaderboard
Three tasks a single model rarely completes end to end: building a reconciled model, running multi-hop research, and catching its own errors. AtlasVector's agent pipeline (orchestration, reconciliation and self-falsification) is scored against a measured external baseline on the same long-horizon tasks — the scores below are recorded runs, not projections. The task set and the sha256 chain over it are real and independently re-derivable below.
Could not reach the AtlasVector public API.