Run configuration
Runs currently supportenvironment: "aws" only. Optional properties sets the AWS region, S3 bucket, log group, and log retention. Managed runs use deployment settings. Runs using your AWS credentials may supply their own properties.
New runs retain these locations for retries, resumes, and artifact reads. Older runs without saved properties use the current settings.
Benchmark-service authentication
The tracker forwards its inbound Descope API key only when the benchmark-service origin matches the hosted origin derived from the benchmark name and tracker configuration. A custom origin never receives that key. Custom services authenticate with explicit headers or a stored service secret instead. See Benchmark authentication.Sandbox scheduling
SANDBOX_QUEUE_ENABLED defaults to false. In that mode, runs create sandboxes directly and requests with an explicit priority are rejected. When the flag is true, the tracker queues runs only when the selected sandbox provider has a managed admission pool. Unmanaged providers continue to use direct execution and also reject explicit priority.
Managed shared pools order waiting tasks by run priority, enqueue time, and task ID. Priority ranges from P0 (highest) to P4 (lowest), and admitted runs default to P3 when no priority is supplied. Per-run concurrency still limits how many tasks from that run may be active.
Users can inspect their organization’s scheduler state with valkyrie queue status or client.scheduler.overview(). Both use:
waiting_next_offset and active_next_offset identify the next page, or return null when exhausted. The corresponding capped flag indicates more rows after that page. Offsets apply after organization filtering and do not change global pool positions or totals. Pages reflect live state; queue changes between calls can repeat or skip entries. Active entries use BUILDING, IN_PROGRESS, or EVALUATING status.
Add include_capacity=true to request best-effort CPU, memory, disk, and aggregate GPU capacity observations for each pool referenced by waiting or active work. Active-only pools have a waiting count of zero. Capacity is returned only when every waiting or active run in that pool uses the same managed sandbox-provider configuration and the reconstructed provider still matches the persisted pool. Each capacity_domains entry keeps one canonical provider target and sandbox class separate; the API never combines distinct domains. GPU-capable providers can also report the allowed GPU types for each domain, but availability is aggregate rather than per type. Allowed types are observational display metadata, not authorization to request a GPU type. Access-key, ambiguous, unsupported, timed-out, or unavailable provider configurations return capacity_domains: null without failing the queue snapshot, while an empty list means the provider successfully reported no domains. Values are observational and not reserved. Provider credentials and configuration identifiers remain internal.
Agent library limits
Agent upload and removal use the same authentication and AWS runtime as listing and downloading. Bundles remain in the sharedagents/ prefix. PUT replaces an existing alias; DELETE reports a missing alias as HTTP 404. Existing storage permissions can deny writes with HTTP 403. These endpoints do not grant additional permissions.
Set these positive-integer environment variables on the Tracker process to override archive limits:
Tracker rejects nonpositive or invalid settings at startup. Exceeded limits return HTTP 413 before storage publication. It checks Content-Length when present and always counts streamed bytes. It checks ZIP metadata before decompressing, then validates actual bytes, integrity, paths, symlinks, and the full agent contract. Temporary disk space must accommodate the upload; validation completes before replacing the stored ZIP.
Run it locally
http://localhost:8000. No .env file is required: Docker Compose reads AWS credentials from your shell environment.
Individual targets:
Tests
tests/integration/live requires services/tracker/.env:

