Testing
Orka has four kinds of tests: Go unit tests, controller integration tests against a real API server (envtest), end-to-end tests against a throwaway Kind cluster, and frontend tests in the browser-like Vitest environment. This page describes what each covers and how to run it.
Running tests
# Run test pipeline (manifests, generate, fmt, vet, then Go tests)
make test
# Run Go tests with coverage report
make test
go tool cover -func=cover.out | grep total
# Run release automation and workspace cleanup tests without a cluster
go test ./cmd/build/release ./scripts/tests
# Run frontend tests
make ui-test # or: cd ui && bun run test
make ui-test-coverage # or: cd ui && bun run test:coverage
# Run E2E tests (requires isolated Kind cluster)
make test-e2e
# Run only the deterministic Gateway live E2E
KIND_CLUSTER=orka-gateway-e2e \
E2E_GATEWAY=true \
E2E_EPHEMERAL_CLUSTER=true \
E2E_GINKGO_FOCUS="Gateway live E2E" \
make test-e2e
# Run Agent Substrate E2E (requires Docker, Go, git, curl, kind, kubectl, ko, jq)
bash scripts/agent-substrate-e2e.sh
# Lint
make lint
make lint-fix
make ui-lint
Local environment notes
- Script test suites need bash >= 4. The suites under
scripts/tests/rely onset -estopping on failed(( ))arithmetic, which macOS's stock bash 3.2 does not honor — failures would pass silently. The suites refuse to run under bash < 4; on macOS install a modern bash (brew install bash) and invoke the suites with it. - A green Gateway E2E can be an empty one. Without
E2E_GATEWAY=truethe gateway specs skip themselves and Ginkgo printsRan 0 of N Specswith aSUCCESSexit. If you see that line, nothing was validated — re-run with the environment shown above. (CI fails this shape explicitly.) - Running
go testdirectly on controller packages needs envtest assets.make testwiresKUBEBUILDER_ASSETSautomatically; for a barego teston packages that start an envtest API server, export it first:KUBEBUILDER_ASSETS="$(bin/setup-envtest use -p path)".
Test structure
Go tests
Tests use Ginkgo + Gomega (BDD style) for controller/integration tests and standard Go testing for unit tests.
| Package | Test Files | Coverage Areas |
|---|---|---|
cmd/build/release/ | workflow_test.go, prepare_test.go, publish_test.go, version_test.go | Release version edits, trusted tooling, branch races, artifact identity, approval evidence, chart publication, and retries. Uses local Git repositories and Helm; GitHub and registry responses are fixtures. |
scripts/tests/ | workspace_lifecycle_test.go | Workspace cancellation, session archival, suspension, and deletion ordering using Bash and jq fixtures. |
internal/api/ | handlers_test.go, internal_handlers_test.go, auth_test.go, middleware_test.go, pagination_test.go, server_test.go, openai_compat_test.go | REST API handlers, internal API handlers, memory/session APIs, authentication, middleware, pagination, OpenAI compatibility |
internal/controller/ | task_controller_test.go, agent_controller_test.go, tool_controller_test.go, session_manager_test.go, job_builder_test.go, repositoryscan_controller_test.go, webhook_test.go | Reconciliation logic, session management, job building, coordination enforcement, repository scan mapper/finding/patch ingestion |
internal/security/ | security_test.go, contracts_test.go | Repository security artifact contracts, v2 evidence validation, fingerprinting, bounded context manifests, prompt helpers |
internal/security/slices/ | mapper_test.go | Deterministic review-slice mapper coverage for Go, Node/TypeScript, Python, workflows, scripts, config, path skipping, and stable output |
internal/store/sqlite/ | security_store_test.go | Repository security records, findings, review slices, dropped finding diagnostics, patch proposals |
internal/store/sqlite/ | schema_test.go, integration_test.go | Complete current schema, incompatible-layout rejection without data loss, repeated reopening with stable record identities and ordering |
internal/llm/ | provider_test.go | Provider registry |
internal/llm/anthropic/ | provider_test.go | Anthropic API integration |
internal/llm/openai/ | provider_test.go | OpenAI API integration |
internal/metrics/ | metrics_test.go | Prometheus metrics recording |
internal/tools/ | registry_test.go, memory tool tests, coordination tool tests, PR tool tests, agent-management tool tests, integration_test.go | Built-in tool implementations, memory tools, coordination tools, PR tools, agent management tools |
internal/worker/ | tool_executor_test.go | Custom Tool CRD executor |
workers/ai/ | main_test.go | AI worker functions |
workers/general/ | main_test.go | General worker functions |
internal/harness/v2/, internal/acp/, workers/acp/ | ACP contract, client, supervisor, and conformance tests | RuntimeSession lifecycle, exact fences, duplicate handling, event bounds, cancellation, workspace deltas, process cleanup, and redaction |
E2E tests
End-to-end tests run against a dedicated Kind cluster:
| Test File | Coverage |
|---|---|
test/e2e/e2e_test.go | Core task lifecycle |
test/e2e/agent_test.go | Agent task execution |
test/e2e/agent_copilot_test.go | Copilot built-in profile admission plus exact digest-pinned RuntimePool image selection, without requiring live provider authentication |
test/e2e/agent_claude_test.go | Claude runtime |
test/e2e/agent_workspace_test.go | Workspace/git clone |
test/e2e/agent_session_test.go | Session continuity |
test/e2e/autonomous_mode_test.go | Autonomous iterations, max-iteration stop, Plan API, suspend behavior |
test/e2e/coordination_advanced_test.go | cancel_task, inter-task messaging, auto-retry, dynamic agent create/delete |
test/e2e/pr_workflow_test.go | PR tool workflow (create_pull_request, review/comment/merge) and workspace PR env wiring |
test/e2e/api_coverage_test.go | Sessions, agent update API, single-tool API, auth validation, secrets API, chat delete, non-autonomous plan 404 |
test/e2e/chat_advanced_test.go | JSON chat mode, agentRef chat routing, management tools via chat |
test/e2e/security_enforcement_test.go | Non-root execution, read-only filesystem, deny-pattern enforcement, kube-system chat block |
test/e2e/agent_advanced_test.go | Skills ConfigMap wiring, agent resource propagation, session maxMessages behavior |
test/e2e/workspace_advanced_test.go | Workspace source/ref/subPath, separate read/publication credential roles, delivery status, and Session behavior |
test/e2e/provider_advanced_test.go | Provider rate-limit config coverage |
test/e2e/live_copilot_proxy_test.go | Native type: ai Provider compatibility against the separately deployed copilot-proxy service; this is separate from built-in Copilot ACP RuntimePool coverage |
test/e2e/live_chat_api_test.go | Live chat SSE and JSON transport/session coverage using a proxy-backed Provider |
test/e2e/live_anthropic_compat_test.go | Live Anthropic-compatible /anthropic/v1/models and /anthropic/v1/messages coverage with default tools-enabled behavior |
test/e2e/live_agent_runtime_matrix_test.go | Historical live Codex/Claude provider execution plus a digest-pinned Copilot image smoke assertion; it is not the canonical ACP release gate |
test/e2e/gateway_test.go | Authenticated Gateway ingress through a deterministic external AgentRuntime, including TLS adapter readiness, invalid bearer rejection, accepted and duplicate events, Task execution, completed events, delivered replies, idempotency, and Task/delivery correlation |
.github/workflows/gateway-e2e.yml | Focused, model-free, secret-free Gateway live E2E in Kind using generated bearer tokens, an ephemeral CA, the TLS reference adapter, and the deterministic echo runtime |
.github/workflows/live-agent-sandbox-e2e.yml / scripts/live-agent-sandbox-e2e.sh | Live upstream agent-sandbox Kind validation for Orka agent workspace claim, sandbox execution, delete cleanup, retained-session reuse, and token scrubbing using a fake model-free Claude runtime |
.github/workflows/repository-monitor-smoke.yml | Focused RepositoryMonitor smoke coverage for store CRUD, API handlers, pull request event handling, targeted single-PR inventory runs, controller queue/review flow, blocked status counts, read-only review task job building, result stdout forwarding, create_pr_monitor repository URL and credential validation, GitHub tool repo_url scope enforcement, and PR review marker tooling |
.github/workflows/security-scan-e2e.yml / scripts/security-scan-e2e.sh | Secret-free repository security scan Kind validation against pinned sozercan/nodejs-goof using the real mapper, deterministic fake Codex analyzer, v2 finding ingestion/drop diagnostics, threat-model rejection, idempotent rescan, and HITL no-auto-patch gating |
test/e2e/tools_test.go | Built-in tools (including web_fetch, file_write) and custom Tool CRD |
test/e2e/scheduled_task_test.go | Cron scheduling, suspend, concurrencyPolicy: Forbid, history-limit cleanup |
test/e2e/task_lifecycle_test.go | Timeout/retry/cancel plus session serialization and lock release |
scripts/agent-runtime-e2e.sh | Canonical deployed-cluster ACP smoke/release gate for Codex, OpenCode, Claude, and Copilot RuntimePools, exact Pod/runtime identity, workspace read/write, continuation/fork, cancellation/timeout, restart/replacement, publication/PR verification, drain/scale-to-zero, immutable images, and cleanup |
.github/workflows/agent-runtime-e2e.yml / scripts/agent-runtime-kind-e2e.sh | Trusted-branch/nightly/manual agent runtime smoke that calls real model providers, requires credentials, and bootstraps an ephemeral Kind cluster, Vekil, and the production ACP topology before invoking the canonical validator |
.github/workflows/release-qualification.yml / scripts/agent-runtime-kind-e2e.sh | Environment-approved chart and runtime acceptance, real local Git publication tests, GitHub API fixture tests, and cleanup; no stored Git publication credentials |
The Gateway Live E2E workflow (.github/workflows/gateway-e2e.yml) runs on manual dispatch and on pull requests or pushes that touch Gateway-relevant source, configuration, E2E, image, or dependency paths. It creates a dedicated Kind cluster, generates disposable TLS and bearer credentials, deploys the TLS reference adapter and deterministic echo AgentRuntime, and verifies invalid bearer rejection, accepted and duplicate ingress, runtime-backed Task completion, final delivery, idempotency, and correlation metadata. The workflow is model-free and secret-free; it does not use repository or provider credentials.
The Repository Monitor Smoke workflow runs in GitHub Actions on pull requests and pushes that touch the workflow, API, controller, CRD/config, worker, or Go dependency paths. It creates the UI embed stub and runs focused go test selections for the monitor store, API handlers, GitHub pull request event handling, targeted single-PR inventory runs, controller queue/review flow, blocked status counts, read-only review job construction, result stdout forwarding, create_pr_monitor repository URL and credential validation, GitHub tool repo_url scope enforcement, and PR review marker signing/detection tooling. The workflow is secret-free: exact PR event queueing is tested with synthetic signed webhook payloads and fake GitHub clients rather than live repository credentials. The normal Go Tests workflow runs make test for non-doc code changes and covers worker-level PR review diff context generation.
Repository security E2E coverage should include initial deterministic slice creation, incremental scan behavior, invalid v2 evidence being dropped and visible through API, validation task persistence, successful verified patch proposals, and patch proposals with missing or mismatched artifacts staying not ready.
E2E key requirements
-
scripts/agent-runtime-e2e.sh --context <context>is the canonical ACP deployed-cluster validator. Its default mode is a smoke test; setRELEASE_GATE=1for full runtime acceptance, including Task result/fork and scale-to-zero scenarios. The Kind wrapper also requires uncached publication tests. A smoke result cannot qualify a release. Live GitHub publication is available only through an explicit local opt-in. -
scripts/agent-runtime-kind-e2e.shis the CI/local bootstrap entrypoint. It creates an ephemeral Kind cluster, deploys Vekil and the production ACP topology with digest-pinned local images, and then calls the canonical script. -
.github/workflows/agent-runtime-e2e.ymlruns on relevant pushes to the default branch, nightly, and manual dispatch. It requires the repository'sCOPILOT_GITHUB_TOKENsecret so Vekil can exercise Codex, OpenCode, Claude, and Copilot against real model providers without mounting provider credentials into RuntimePools. The workflow rejects dispatches from other branches before using the secret. It runs unattended as ordinary CI without a deployment environment. When migrating fromlive-acp-runtime-smoke, ensure the provider secret is available at repository scope. After the renamed workflow succeeds, retire the old environment and preserve its deployment history. -
.github/workflows/release-qualification.ymlusesworkflow_dispatchand is serialized. The release workflow dispatches it automatically. Restrict therelease-qualificationenvironment tomainand exact permitted release branches, require a trusted reviewer, and disable administrator bypass before exposing model-provider credentials. It accepts a full source SHA that must equal both the dispatched workflow commit and selected branch head. GitHub observations use the job's temporaryGITHUB_TOKENwithcontents: read.COPILOT_GITHUB_TOKENsupplies model-provider access and can come from the existing repository secret. No stored Git publication token or fork setting is needed. Publication tests use real temporary Git repositories and local GitHub API test servers. They do not prove the complete cluster-to-GitHub publication path or live organization permissions. Release qualification documents trusted dispatch, required test evidence, report verification, and cleanup. Qualification needs the report for the exact candidate; smoke success is insufficient. -
Neither
Agent Runtime E2EnorRelease Qualificationruns forpull_request, so PR-controlled code never receives provider credentials. Both check out with persisted credentials disabled and expose secrets only to the final local script step. The release gate addsactions: readtocontents: readto verify the environment branch rules and download candidate artifacts using the native token. Neither workflow requestsid-token: write; GitHub OIDC validation remains isolated inlive-github-oidc-e2e.yml. -
The release gate additionally requires digest-pinned controller, Publisher, Codex, OpenCode, Claude, and Copilot images; a Ready central provider proxy; the
config/acp-productionVekil ingress boundary; durable controller/Publisher PVCs; a distinct publication fork; and authenticatedghaccess for independent remote and PR verification/cleanup. -
Canonical ACP acceptance runs all Codex scenarios, including publication, before deleting the Codex pool and starting OpenCode. It validates OpenCode native ACP read, continuation, and read-intent tool policy, deletes the OpenCode pool, starts Claude, and then runs live Copilot read/continuation after Claude cleanup. Every provider phase verifies exact Pod UID, image ID, runtime instance/profile, and RuntimeSession identity. Codex additionally covers cancellation and timeout settlement, controller restart, pool replacement, drain/scale-to-zero, and publication/remote delivery; release mode also exercises Task forks.
-
E2E_OPENAI_API_KEYandE2E_ANTHROPIC_API_KEYremain inputs for older nativetype: aitest cases. They are not mounted into built-in ACP RuntimePools. -
COPILOT_GITHUB_TOKENalso remains the credential forlive-copilot-proxy-e2e.yml, which covers the external proxy as native Provider test infrastructure.Agent Runtime E2Eprovides the provider-execution evidence for the built-in RuntimePool profiles. -
Structural e2e tests for native worker Jobs run without external model keys.
-
Live Agent Sandbox E2EandAgent Substrate E2Edo run workspace-backed ACP Tasks end to end against a local model fixture, but they are not the full release gate: external-provider execution and clean-room publication stay withAgent Runtime E2EandRelease Qualification. Every Substrate run includes DataOnly suspension, cold continuation, and checkpoint file recovery against the unmodified upstream provider. -
Security Scan E2E is secret-free and model-free, but requires Docker plus the local Go, Kind, kubectl, curl, and jq toolchain.
-
Gateway Live E2E is model-free and secret-free. Its focused invocation sets
E2E_GATEWAY=trueandE2E_EPHEMERAL_CLUSTER=true; the last flag skips per-resource suite cleanup because the caller deletes the entire Kind cluster. -
GitHub Actions
id-token: writepermission: required by the live GitHub OIDC workflow. For local/manual runs ofscripts/live-github-oidc-e2e.sh, setORKA_GITHUB_OIDC_TOKENto a valid JWT instead. Provider-specific transaction-token E2E lives in the external integration repositories. -
E2E_LIVE_COPILOT_PROXY_BASE_URL(orE2E_COPILOT_PROXY_BASE_URL/COPILOT_PROXY_BASE_URL): enables the focused live copilot-proxy spec against a running proxy -
E2E_LIVE_COPILOT_PROXY_SERVICE_NAMESPACE,E2E_LIVE_COPILOT_PROXY_SERVICE_NAME,E2E_LIVE_COPILOT_PROXY_SERVICE_PORT: direct-access proxy coordinates used by the legacy Provider, Chat, and Anthropic compatibility specs. -
E2E_LIVE_ACP_PROVIDER_PROXY_SERVICE_NAMESPACE,E2E_LIVE_ACP_PROVIDER_PROXY_SERVICE_NAME,E2E_LIVE_ACP_PROVIDER_PROXY_SERVICE_PORT: model-discovery coordinates for the ACP RuntimePool matrix. The full live script pins these tovekil-system,vekil, and1337so built-in RuntimePools traverse the production provider-proxy DNS and NetworkPolicy boundary.
Live provider/chat tests prefer gpt-5-mini and claude-haiku-4.5 when available.
The runtime smoke and release qualification scripts default to gpt-5.4-mini
for Codex and OpenCode, claude-haiku-4.5 for Claude, and gpt-5.3-codex for
Copilot. Override these with ACP_E2E_CODEX_MODEL, ACP_E2E_OPENCODE_MODEL,
ACP_E2E_CLAUDE_MODEL, or ACP_E2E_COPILOT_MODEL. Configured models must pass
the Vekil catalog and streaming probes.
Run the trusted smoke bootstrap locally with the token exported in the shell rather than placed on a command line:
read -rsp 'Copilot provider token: ' COPILOT_GITHUB_TOKEN && echo
export COPILOT_GITHUB_TOKEN
export ACP_E2E_OPENCODE_MODEL=openai/gpt-5.4-mini
export ACP_E2E_OPENCODE_CONTEXT_WINDOW=32768
export ACP_E2E_OPENCODE_MAX_TOKENS=4096
ACP_E2E_KIND_TAG=local bash scripts/agent-runtime-kind-e2e.sh
For a local qualification diagnostic, use read-only GitHub CLI access and the same model-provider configuration as the smoke run. GitHub's automatic job token is available only inside Actions. To use that token without configuring local GitHub access, dispatch Release Qualification.
The local wrapper runs the required Git and PR fixture tests before creating the cluster:
export RELEASE_GATE=1
export ACP_E2E_WRITE_CREATE_PR=0
export ACP_E2E_REPO=https://github.com/orka-agents/orka.git
export ACP_E2E_REF="$(git rev-parse HEAD)"
export ACP_E2E_BASE_BRANCH=main
bash scripts/agent-runtime-kind-e2e.sh
The fixture tests, image build, and runtime phases use the same candidate.
The wrapper requires a clean checkout at that commit, and the candidate must
equal the selected branch head. Local reports are saved under
bin/acp-release-*/acceptance.json, or
ACP_E2E_REPORT_FILE when set. --keep-cluster leaves cleanup pending and cannot
qualify a release; use the trusted workflow and report verifier above for release
records.
The existing local live GitHub diagnostic remains available with
ACP_E2E_WRITE_CREATE_PR=1 and separately authorized canary credentials. Its
four-role and distinct-fork requirements are listed in the validator's --help.
The hosted release workflow cannot enable it and explicitly reports live GitHub
publication as not_tested.
Do not enable shell xtrace for either invocation. The scripts create Kubernetes Secrets without printing their values and redact provider/GitHub token patterns from failure diagnostics.
The live agent-sandbox workflow validates both the direct workspace-adapter lifecycle and the initial workspace-backed Orka harness v2 happy path. It builds the real Codex supervisor, routes a prompt through a local Responses-compatible fixture, waits for the Task to succeed, verifies provider-neutral status, and cleans up the dedicated RuntimePool. It does not replace release qualification or provide publication evidence. Its direct-adapter assertions also include:
- the adapter creates a v1beta1
SandboxClaimwith the expectedwarmPoolRefand executes a command with caller-supplied env inside the sandbox cleanupPolicy: deleteremoves the generatedSandboxClaimcleanupPolicy: retainplusreusePolicy: sessionreattaches to the deterministic session claim- retained workspace state persists across tasks
The RepositoryMonitor validation bundle covers the canonical signed orka:implement entrypoint, durable intake, pause controls, and replay against GitHub fixtures.
The live GitHub OIDC workflow (.github/workflows/live-github-oidc-e2e.yml) runs scripts/live-github-oidc-e2e.sh in GitHub Actions with id-token: write. It builds the controller from the PR, deploys it to a fresh Kind cluster, configures the GitHub OIDC issuer and workflow audience, fetches a real Actions OIDC token, and validates:
- unauthenticated API requests return
401 - OIDC-authenticated Task creation returns
201 - the created Task contains verified
spec.requestedByprovenance - top-level
requestedByand nestedspec.requestedByclient tampering are rejected with400 - the OIDC token does not appear in controller logs
The Agent Substrate workflow runs the official pin in
hack/agent-substrate/upstream.env without provider source patches. Every run
includes direct sealed execution and files, MCP, ACP, controller restart,
DataOnly suspension, cold continuation, checkpoint export and restore,
cancellation, timeout, and cleanup. Native unit tests cover TLS and credential
rotation, lost responses, source identity changes, reference races, and explicit
recovery. Tests do not supply fork-only lifecycle preconditions.
The ACP file scenario uses the real Codex runtime with a deterministic Responses
fixture. It writes a file, exports a checkpoint, changes and deletes the source
workspace, then restores the checkpoint and checks the original bytes through a
shell read. The fixture requires successful tool output from the current turn.
The file Tasks retain read-only intent, so the suite requires successful
execution and ReadOnlyWorkspaceModified delivery rejection. The workflow also
runs a Linux regression as root to verify durable file ownership after session
deletion, drain, and supervisor shutdown.
The lifetime case starts a real shell command with a 300-second hold in a
workspace with a 120-second maxLifetime. It verifies that expiry cancels the
original prompt and removes its worker before the command can finish. The Task
has a longer timeout, and the test does not cancel or delete it to force expiry.
Session, workspace, pool, and saved-data cleanup must finish afterward.
bash scripts/tests/agent-substrate-e2e-hardening-test.sh
KEEP_CLUSTER=1 bash scripts/agent-substrate-e2e.sh
Use a new dedicated KIND_CLUSTER or explicitly select
SUBSTRATE_REUSE_CLUSTER=1. The installer never recreates an existing cluster.
The provider owns its gVisor kind setup; the kubeconfig stays under the run's
bin/ directory. Live conformance requires a working Docker engine.
Frontend tests
Frontend tests use Vitest + Testing Library + MSW.
cd ui && bun run test # what CI runs
cd ui && bun run test:coverage # adds the coverage report and threshold check
ui/vite.config.ts declares thresholds of 95% statements, 80% branches, 90% functions, and
95% lines. Those only apply to bun run test:coverage. The ui-test job in
.github/workflows/test.yml runs plain bun run test, so a pull request that drops coverage
still passes CI. Run the coverage command yourself before assuming a number.
Testing patterns
Table-driven tests
tests := []struct {
name string
input string
want string
wantErr bool
}{
{"valid", "input", "output", false},
{"invalid", "bad", "", true},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
// test logic
})
}
Fake Kubernetes client
scheme := runtime.NewScheme()
corev1alpha1.AddToScheme(scheme)
corev1.AddToScheme(scheme)
client := fake.NewClientBuilder().WithScheme(scheme).WithObjects(objs...).Build()
HTTP mocking
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.WriteHeader(http.StatusOK)
w.Write([]byte(`{"result": "ok"}`))
}))
defer server.Close()
Fiber test app
app := fiber.New()
app.Get("/test", handler)
req := httptest.NewRequest(http.MethodGet, "/test", nil)
resp, _ := app.Test(req)
Frontend test mocking
// Mock zustand persist middleware
vi.mock('zustand/middleware', () => ({ persist: (fn: unknown) => fn }))
// Use test utils with QueryClient wrapper
import { render } from '@/test/test-utils'
Testing with chat
When testing features via the chat endpoint, use natural prompts — the kind a human would actually type. Never reference internal concepts like agent names, tool names, or implementation details. Describe what you want done, not how the system should do it. The chat should infer the right agents, tools, delegation patterns, and cancellation logic on its own.
Good examples:
- "Research the benefits of Kubernetes and write a technical guide based on the findings."
- "What's the best container orchestration tool? Get me an answer as fast as possible."
- "Draft an outline for a blog post about containers and turn it into a full post."
- "Compare microservices vs monoliths from three angles, then synthesize into a recommendation."
Bad examples:
- "Create a coordinator agent and a researcher agent, then delegate two tasks..."
- "Use the send_message tool to send a message to task msg-receiver..."
- "Have three researchers race to answer..." (users don't think in terms of "researchers")
- "Use the first answer and cancel the others." (the system should infer this automatically)