> ## Documentation Index
> Fetch the complete documentation index at: https://docs.valkyrie.vals.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Tracker service

> Architecture, local commands, and test suites for the tracker backend.

The tracker is the FastAPI backend. It serves the CLI and SDK API, tracks run and task state in PostgreSQL, writes artifacts to S3, and publishes task dispatches through Redis. Each dispatch names an executor release. ExecutorHost picks the dispatch up and runs that release in a sandbox.

Local Docker Compose starts the tracker, PostgreSQL, and Redis only. It does not run ExecutorHost or create an active executor release, so a locally started run never executes. Use a deployed environment to run benchmarks.

## Run configuration

Runs currently support `environment: "aws"` only. Optional `properties` sets the AWS region, S3 bucket, log group, and log retention. Managed runs use deployment settings. Runs using your AWS credentials may supply their own properties.

New runs retain these locations for retries, resumes, and artifact reads. Older runs without saved properties use the current settings.

## Benchmark-service authentication

The tracker forwards its inbound Descope API key only when the benchmark-service origin matches the hosted origin derived from the benchmark name and tracker configuration. A custom origin never receives that key. Custom services authenticate with explicit headers or a stored service secret instead. See [Benchmark authentication](/benchmarks/authentication).

## Sandbox scheduling

`SANDBOX_QUEUE_ENABLED` defaults to `false`. In that mode, runs create sandboxes directly and requests with an explicit priority are rejected. When the flag is `true`, the tracker queues runs only when the selected sandbox provider has a managed admission pool. Unmanaged providers continue to use direct execution and also reject explicit priority.

Managed shared pools order waiting tasks by run priority, enqueue time, and task ID. Priority ranges from P0 (highest) to P4 (lowest), and admitted runs default to P3 when no priority is supplied. Per-run `concurrency` still limits how many tasks from that run may be active.

Users can inspect their organization's scheduler state with [`valkyrie queue status`](/reference/cli/queue#status) or [`client.scheduler.overview()`](/reference/sdk/scheduler#overview). Both use:

```text theme={null}
GET /scheduler/overview?waiting_limit=100&active_limit=100&waiting_offset=0&active_offset=0
```

The endpoint uses the request's resolved organization and never returns rows from another organization. Its summary contains total waiting, building, in-progress, and evaluating counts. It also returns per-pool waiting counts plus bounded waiting and active entries. Each entry limit defaults to 100 and accepts 1–200. Independent nonnegative offsets default to zero. `waiting_next_offset` and `active_next_offset` identify the next page, or return `null` when exhausted. The corresponding capped flag indicates more rows after that page. Offsets apply after organization filtering and do not change global pool positions or totals. Pages reflect live state; queue changes between calls can repeat or skip entries. Active entries use `BUILDING`, `IN_PROGRESS`, or `EVALUATING` status.

Add `include_capacity=true` to request best-effort CPU, memory, disk, and aggregate GPU capacity observations for each pool referenced by waiting or active work. Active-only pools have a waiting count of zero. Capacity is returned only when every waiting or active run in that pool uses the same managed sandbox-provider configuration and the reconstructed provider still matches the persisted pool. Each `capacity_domains` entry keeps one canonical provider target and sandbox class separate; the API never combines distinct domains. GPU-capable providers can also report the allowed GPU types for each domain, but availability is aggregate rather than per type. Allowed types are observational display metadata, not authorization to request a GPU type. Access-key, ambiguous, unsupported, timed-out, or unavailable provider configurations return `capacity_domains: null` without failing the queue snapshot, while an empty list means the provider successfully reported no domains. Values are observational and not reserved. Provider credentials and configuration identifiers remain internal.

## Agent library limits

Agent upload and removal use the same authentication and AWS runtime as listing and downloading. Bundles remain in the shared `agents/` prefix. PUT replaces an existing alias; DELETE reports a missing alias as HTTP 404. Existing storage permissions can deny writes with HTTP 403. These endpoints do not grant additional permissions.

Set these positive-integer environment variables on the Tracker process to override archive limits:

| Variable                           | Default              | Bound                                  |
| ---------------------------------- | -------------------- | -------------------------------------- |
| `AGENT_UPLOAD_MAX_BYTES`           | `1073741824` (1 GiB) | Actual uploaded ZIP bytes              |
| `AGENT_ARCHIVE_MAX_EXPANDED_BYTES` | `5368709120` (5 GiB) | Declared and actual decompressed bytes |
| `AGENT_ARCHIVE_MAX_ENTRIES`        | `100000`             | ZIP entries, including directories     |

Tracker rejects nonpositive or invalid settings at startup. Exceeded limits return HTTP 413 before storage publication. It checks Content-Length when present and always counts streamed bytes. It checks ZIP metadata before decompressing, then validates actual bytes, integrity, paths, symlinks, and the full agent contract. Temporary disk space must accommodate the upload; validation completes before replacing the stored ZIP.

## Run it locally

```bash theme={null}
make tracker-service
```

This builds and starts the local API and infrastructure, then tails the tracker logs. The API is available at `http://localhost:8000`. No `.env` file is required: Docker Compose reads AWS credentials from your shell environment.

Individual targets:

```bash theme={null}
make build    # Build the tracker image
make run      # Start the tracker, PostgreSQL, and Redis
make stop     # Stop the local stack
make clean    # Stop the stack and remove local images
make logs     # Tail tracker logs
```

## Tests

```bash theme={null}
make test                    # Unit and local integration tests, 85% total coverage
make test-unit               # Unit tests and Alembic migrations
make test-alembic            # Alembic migration tests only
make test-integration-local  # Local API and Postgres tests
make test-integration-live   # AWS, benchmark service, and sandbox tests
```

`tests/integration/live` requires `services/tracker/.env`:

```env theme={null}
# AWS: uses Secrets Manager and performs live S3 and CloudWatch create/read/delete tests
AWS_ACCESS_KEY_ID=
AWS_SECRET_ACCESS_KEY=
AWS_DEFAULT_REGION=us-east-1
AWS_SESSION_TOKEN=              # Optional, required for temporary credentials

# Test infrastructure
TEST_AWS_S3_BUCKET=             # S3 bucket for agent artifacts
TEST_LOG_GROUP=                 # CloudWatch log group
TEST_DAYTONA_SECRET_NAME=       # Secrets Manager secret holding the Daytona provider config

# Benchmark service
BENCHMARK_SERVICE_BASE_URL=     # Domain of the benchmark service
BENCHMARK_SERVICE_CLOUDMAP_NAMESPACE=local  # Cloud Map namespace fallback when no base URL is set
BENCHMARK_SERVICE_AUTH_KEY=     # Access key for authenticating with a benchmark service
```

## Migrations

```bash theme={null}
make migrate-gen   # Generate a migration from model changes
```

See [Database and migrations](/contributing/database) for the full guide.

### Runtime configuration rollout

Stop admission and drain active runs before deploying Tracker with a compatible executor release. Verify a new run and a stopped-run resume before reopening admission.
