Skip to main content

Configuration

Custom resources​

Task​

The core work unit. Supports container commands, native AI prompts, or Orka harness v2 coding-agent RuntimeSessions.

apiVersion: core.orka.ai/v1alpha1
kind: Task
metadata:
name: my-task
spec:
type: ai # or "container" or "agent"
agentRef:
name: my-agent
prompt: "Analyze the latest Kubernetes security best practices"
execution:
runtimeClassName: gvisor
nodeSelector:
sandbox-runtime: gvisor
sessionRef:
name: my-session
create: false # default: false
append: true
maxMessages: 50
priority: 500
timeout: 5m
retryPolicy:
maxRetries: 3
backoffMultiplier: 2
initialDelay: 10s
webhookURL: "https://example.com/webhook"
# Scheduled/recurring task fields (optional)
schedule: "0 */6 * * *" # Cron expression
timeZone: "America/New_York" # IANA timezone
concurrencyPolicy: Forbid # Allow or Forbid concurrent runs
startingDeadlineSeconds: 100 # Deadline for starting missed scheduled runs (default: 100)
suspend: false
successfulRunsHistoryLimit: 3
failedRunsHistoryLimit: 1

For type: agent, repository and delivery policy belong at top-level spec.workspace:

workspace:
intent: write
gitRepo: https://github.com/example/project.git
branch: main
readCredentialRef:
name: project-source-read
publicationGitRepo: https://github.com/example/project.git
publicationReadCredentialRef:
name: project-target-read
publicationCredentialRef:
name: project-target-write
forgeCredentialRef:
name: project-forge
pushBranch: orka/example-change
prBaseBranch: main
createPR: true

readCredentialRef is source-read only; publicationReadCredentialRef is target preflight/verification only; publicationCredentialRef is target-write only; and forgeCredentialRef is PR-reconciliation only. The controller freezes each selected Secret version and the credential broker releases it only to the Workspace/Publisher for the exact operation. None enters the ACP runtime process tree.

Agent​

Reusable agent configurations with model settings, tools, skills, and optional agent-to-agent coordination.

apiVersion: core.orka.ai/v1alpha1
kind: Agent
metadata:
name: researcher-agent
spec:
providerRef:
name: anthropic-prod
execution:
runtimeClassName: kata-qemu
nodeSelector:
sandbox-runtime: kata
model:
temperature: 0.7
maxTokens: 4096
systemPrompt:
inline: "You are a research specialist..."
tools:
- name: web-search
- name: github-search
skills:
- name: skill-researcher
session:
maxMessages: 50
coordination:
enabled: true
autonomous: true # Enable autonomous loop mode
maxIterations: 20 # Max loop iterations (0 = unlimited)
allowedAgents:
- name: coder-agent
maxConcurrentChildren: 5
maxDepth: 3

Coordination fields:

FieldTypeDefaultDescription
enabledboolfalseEnable agent-to-agent coordination tools
autonomousboolfalseEnables autonomous loop mode. When true, the controller re-creates Jobs in a loop instead of marking the task as Succeeded
maxIterationsint320Limits the number of autonomous loop iterations. Only used when autonomous is true. 0 means unlimited
approvalRequiredToolslist[]Custom Tool CRD names that require human approval before execution in enabled autonomous coordination mode. Built-in tools, including request_approval, are rejected
allowedAgentslist[]List of agent names this agent is allowed to delegate to
maxConcurrentChildrenint325Maximum number of concurrent child tasks
maxDepthint323Maximum delegation depth

Auto-injected coordination tools (when enabled: true):

delegate_task, wait_for_tasks, create_container_task, cancel_task, send_message, check_messages, recall_memory, remember, propose_memory, search_transcript, create_pull_request, list_pull_requests, check_pr_review_marker, check_pull_request_ci, merge_pull_request, auto_merge_pull_request, review_pull_request, post_review_comment, create_agent, delete_agent, update_plan

When autonomous: true, request_approval is also injected so the worker can park the task after an explicit human approval request.

Opt-in coordination tools (require explicit spec.tools[] entries on the Agent):

list_issues, get_issue, comment_on_issue

PR review marker environment:

Prompt-orchestrated PR monitors use check_pr_review_marker to produce and detect hidden review markers. These variables are read by the worker Task that runs the tool:

Environment variableDescription
ORKA_PR_REVIEW_MARKER_SECRETOptional stable HMAC key for PR review marker signatures. Use a Kubernetes Secret or another secret injection path.
ORKA_PR_REVIEW_MARKER_PREVIOUS_SECRETSOptional comma-separated previous marker keys accepted during rotation.
ORKA_PR_REVIEW_MARKER_TRUSTED_AUTHOROptional GitHub login trusted for legacy marker compatibility. When omitted, Orka resolves the authenticated GitHub user for the Task credential.

RepositoryScan​

Repository security scan configuration. A RepositoryScan is namespace-scoped and tells Orka which repository to scan, how to schedule incremental scans, and which Agents should perform analysis and remediation.

apiVersion: core.orka.ai/v1alpha1
kind: RepositoryScan
metadata:
name: example-repo
namespace: default
spec:
provider: github
repoURL: "https://github.com/example/app"
owner: example
repository: app
branch: main
ref: "v1.2.3" # optional tag, branch, or commit SHA checkout override
subPath: "services/api" # optional monorepo scope
gitSecretRef: # optional for private repositories
name: github-credentials
forkRepo: "https://github.com/example/app-security-fork" # optional remediation fork
prBaseBranch: main # optional PR base branch override
schedule: "0 2 * * *" # optional cron expression for incremental scans
timeZone: "UTC" # optional IANA time zone
historyDays: 30 # optional initial history window
validationMode: light # off, light, or full
validationMaxFindingsPerRun: 8 # optional auto-validation task cap for light mode
validationMinSeverity: medium # optional auto-validation severity threshold
validationMinConfidence: medium # optional auto-validation confidence threshold
customScanInstructionsRef: # optional ConfigMap-backed additive scanner instructions
name: repo-security-policy
key: policy # optional; defaults to policy
falsePositivePolicyRef: # optional ConfigMap-backed false-positive policy
name: repo-security-policy
key: false-positives
analysisAgentRef:
name: security-reviewer
patchAgentRef: # optional; defaults to the analysis agent when omitted
name: security-patcher
maxFindingsPerRun: 25
suspend: false # pause scheduled incremental scans when true

Spec fields:

FieldTypeRequiredDescription
providerstringNoSource control provider. github is the supported v1 provider and default.
repoURLstringYesRepository URL to scan.
ownerstringNoRepository owner or organization. Inferred from repoURL when omitted.
repositorystringNoRepository name. Inferred from repoURL when omitted.
branchstringNoBase branch to scan. Defaults to the literal main when omitted (not resolved from the repository's actual default branch). Set this explicitly for repositories whose default branch is not main.
refstringNoSpecific git ref, tag, or commit SHA to check out for scan tasks. When ref is set and branch is omitted, scan workspaces check out the ref directly instead of forcing main; PR remediation still uses prBaseBranch or main unless branch is set.
subPathstringNoOptional subdirectory to scan in a monorepo.
gitSecretRefLocalObjectReferenceNoCompatibility source-read Secret for scan Tasks; readCredentialRef takes precedence when both are set. Never used for publication.
readCredentialRefLocalObjectReferenceFor patchesSource clone/read Secret. Required, together with the three publication roles below, before any patch proposal or remediation PR.
publicationReadCredentialRefLocalObjectReferenceFor patchesTarget-repository read Secret used only for publication preflight and independent verification.
publicationCredentialRefLocalObjectReferenceFor patchesTarget-repository write Secret used only for the exact compare-and-swap branch push.
forgeCredentialRefLocalObjectReferenceFor patchesForge API Secret used for controller-side forge reads and PR upkeep: fetching the published remediation commit to derive and verify patch evidence, and reconciling/decorating the remediation pull request. Never mounted into any Task. The four patch roles must reference pairwise-distinct Secrets.
forkRepostringNoWritable fork repository URL for patch proposal branches and remediation PRs.
prBaseBranchstringNoPull request base branch for remediation. Defaults to branch when omitted.
schedulestringNoCron expression for scheduled incremental scans.
timeZonestringNoIANA time zone used by schedule.
historyDaysint32NoHow far back the initial scan should inspect repository history.
validationModestringNoValidation aggressiveness: off, light, or full. Defaults to light.
validationMaxFindingsPerRunint32NoMaximum automatic validation tasks to enqueue per scan run in light mode.
validationMinSeveritystringNoMinimum severity eligible for automatic validation. Defaults are mode-dependent.
validationMinConfidencestringNoMinimum confidence eligible for automatic validation. Defaults are mode-dependent.
customScanInstructionsRefPolicyConfigMapKeyRefNoSame-namespace ConfigMap key containing additive scanner instructions. The ConfigMap must opt in with orka.ai/security-policy: "true" as a label or annotation.
falsePositivePolicyRefPolicyConfigMapKeyRefNoSame-namespace ConfigMap key containing additive false-positive policy text. The ConfigMap must opt in with orka.ai/security-policy: "true" as a label or annotation.
analysisAgentRefAgentReferenceYesAgent used for repository scan runs and threat model generation.
patchAgentRefAgentReferenceNoAgent used for patch proposal runs.
maxFindingsPerRunint32NoBounds accepted scan findings per run after validation and deterministic dropped-finding filters.
suspendboolNoPauses scheduled incremental scans while preserving the scan configuration.

PolicyConfigMapKeyRef uses name plus optional key; when key is omitted, Orka reads the policy key. Policy ConfigMap values are capped at 32 KiB and rejected if they look like they contain secrets, tokens, private keys, or credentials. Custom policy text is additive only; it cannot disable Orka's default evidence, no-secret, or finding-quality rules.

Status fields:

status.phase, status.lastScanID, status.lastScanTaskName, status.lastSuccessfulScanAt, status.lastObservedHeadSHA, status.lastProcessedCommit, status.threatModelVersion, status.findingCounts, and status.conditions summarize the latest scan lifecycle and open findings. Dynamic scan runs, threat models, findings, and patch proposals are stored by the controller and surfaced through the security API/UI rather than embedded directly in the CRD status.

RepositoryMonitor​

Durable GitHub pull request monitor configuration. A RepositoryMonitor is namespace-scoped and tells Orka which repository and branch to inspect, which Claude runtime Agent should review selected PR heads, and which labels or scheduling rules should control review selection.

apiVersion: core.orka.ai/v1alpha1
kind: RepositoryMonitor
metadata:
name: example-app
namespace: default
spec:
provider: github
repoURL: "https://github.com/example/app"
owner: example # optional; inferred from repoURL
repository: app # optional; inferred from repoURL
branch: main
gitSecretRef: # optional for private repositories or higher API rate limits
name: repo-monitor-github
schedule: "*/30 * * * *" # optional cron expression
timeZone: "UTC" # optional IANA time zone
suspend: false
targets:
pullRequests:
enabled: true
includeDrafts: false
maxPerRun: 20
agents:
reviewer:
name: repo-reviewer
review:
event: COMMENT # legacy task input only; does not publish to GitHub
staleReviewTTL: 24h
exactEventEnabled: true # queue exact-head runs from signed PR webhooks
publish:
enabled: true # default false; controller-owned GitHub side effect
mode: summary_with_inline_findings
event: COMMENT # V1 only supports neutral COMMENT reviews
postPassed: false
postNeedsChanges: true
postNeedsHuman: true
postSecuritySensitive: false
sameHeadPolicy: skip
inline:
enabled: true
minPriority: P2
maxComments: 10
policy:
protectedLabels:
- security-sensitive
pauseLabels:
- orka:pause
validation:
image: ghcr.io/example/app-validation@sha256:aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa # must contain /bin/sh and the repository's validation tools

Spec fields:

FieldTypeRequiredDescription
providerstringNoSource control provider. github is the supported v1 provider and default.
repoURLstringYesCredential-free GitHub repository root URL to monitor, such as https://github.com/owner/repo, https://github.com/owner/repo.git, or git@github.com:owner/repo.git. Pull request, issue, branch/tree, blob/file, commit, query-string, fragment, non-GitHub, HTTP, and embedded-credential URLs are rejected.
ownerstringNoRepository owner or organization. Inferred from repoURL when omitted.
repositorystringNoRepository name. Inferred from repoURL when omitted.
branchstringNoBase branch used for pull request inventory. Defaults to main.
gitSecretRefLocalObjectReferenceNoGit Secret containing token, password, or GITHUB_TOKEN for GitHub API access and same-repository PR checkout. This is separate from the reviewer Agent's runtime credential Secret.
schedulestringNoCron expression for scheduled monitor runs.
timeZonestringNoIANA time zone used by schedule.
suspendboolNoPauses scheduled monitor runs while preserving the monitor configuration.
targets.pullRequests.enabledboolNoEnables pull request monitoring. Currently this must be true or omitted.
targets.pullRequests.includeDraftsboolNoSelect draft pull requests for review when true. Defaults to false.
targets.pullRequests.maxPerRunint32NoMaximum selected PRs per run. Defaults to 20; allowed range is 1 to 100.
agents.reviewerAgentReferenceYesClaude runtime Agent used for read-only PR review tasks. The Agent must reference a Secret in the monitor namespace with ANTHROPIC_API_KEY or ANTHROPIC_FOUNDRY_API_KEY.
review.eventstringNoLegacy/default review event value included in review task input. It does not publish to GitHub; use review.publish.event. Defaults to COMMENT.
review.publish.enabledboolNoEnables controller-owned GitHub pull request review publishing. Defaults to false.
review.publish.modestringNosummary_only or summary_with_inline_findings. Inline comments are only attempted for changed RIGHT-side diff lines.
review.publish.eventstringNoGitHub review event submitted by the controller. V1 only supports neutral COMMENT reviews; APPROVE and REQUEST_CHANGES are rejected.
review.publish.postPassedboolNoPost clean/passed reviews when true. Defaults to false.
review.publish.postNeedsChangesboolNoPost needs_changes reviews when true. Defaults to true.
review.publish.postNeedsHumanboolNoPost needs_human reviews when true. Defaults to true.
review.publish.postSecuritySensitiveboolNoAllow public publishing of security_sensitive results. Defaults to false; sensitive findings are skipped by default.
review.publish.sameHeadPolicystringNoDuplicate policy for the same monitor, PR, and head SHA. V1 only supports skip.
review.publish.inline.enabledboolNoEnables inline GitHub review comments when mode is summary_with_inline_findings.
review.publish.inline.minPrioritystringNoLowest priority eligible for inline comments (P0-P3). Defaults to P2; lower-priority findings remain in the summary.
review.publish.inline.maxCommentsint32NoMax inline comments per GitHub review. Defaults to 10, allowed range 0 to 50.
review.staleReviewTTLdurationNoRe-review an unchanged head after the previous accepted review is older than this duration.
review.exactEventEnabledboolNoQueue exact-head monitor runs from signed GitHub pull request webhook events when true.
policy.protectedLabelslistNoPR labels that block automated review selection.
policy.pauseLabelslistNoPR labels that pause monitor automation for that item.
validation.imagestringNoContainer image for isolated pull request validation. The reviewer inspects the checked-out code and selects one shell command; Orka fixes the image and exact read-only PR head. The image must contain /bin/sh and every tool the repository may require.

When validation.image is set, the reviewer must call run_validation once and wait for its child Task to finish before returning a passed verdict. A missing, failed, malformed, or stale validation Task blocks passed and merge-ready state. Orka records the image and status with the review and retains only a SHA-256 digest of the selected command in its durable validation binding. Validation stdout and stderr are suppressed so repository or fixture secrets cannot enter Pod logs or result storage. The command runs against a read-only checkout with deny-all ingress and egress, so the image must already contain every required tool and dependency. If a Go repository needs golangci-lint, for example, use a Go-based image that includes dependencies and golangci-lint. The same rule applies to offline Terraform, Azure CLI, or other repository-specific checks. Maintainers normally set this image once per RepositoryMonitor; commands, args, credentials, and network access are not part of the monitor configuration.

Before upgrading a monitor that uses the former validation.mode and validation.commands fields, replace them with validation.image. The CRD retains the old fields so the controller can detect existing policies, but it does not execute them. A non-empty legacy command list puts the monitor in Error with reason LegacyValidationCommandsUnsupported; Orka does not silently disable validation. Apply the updated CRD before upgrading the controller, then update each affected monitor with a digest-pinned image and remove the legacy fields.

targets.issues, the orka:implement label command, issue triage/research/planning/implementation, PR review/repair, review.requireGreenCI, and exact-head readiness are active RepositoryMonitor workflows. targets.commits remain rejected until commit inventory is implemented. Review tasks check out the exact PR head and receive generated read-only context files under /workspace/.git/orka/: pr-review.md, pr-review.files, and pr-review.diff. GitHub publishing, branch pushes, PR creation, label consumption, and readiness publication are controller-owned and audited through mutation records; read-only agents never receive the GitHub mutation token. GitHub's per-PR auto-merge setting owns merging. Orka never enables auto-merge or calls the merge endpoint; when it is disabled, the PR stays open and ready.

Status fields:

status.phase, status.lastRunID, status.lastRunTime, status.lastSuccessfulRunTime, status.observedGeneration, status.openPullRequests, status.pendingReviews, status.activeRepairs, status.blockedItems, status.mergeReadyItems, and status.conditions summarize the monitor lifecycle and current queue. Dynamic runs, PR items, review records, repair records, command events, and audit events are stored by the controller and surfaced through the monitor API/UI rather than embedded directly in CRD status.

See Repository Monitors for the workflow, API examples, and current limits.

Execution​

Native ai and container Tasks support spec.execution for worker Pod runtime selection and placement. Built-in ACP agent Tasks do not accept per-Task placement or custom resources; reviewed RuntimePool profiles own those settings. They can request an execution workspace through spec.execution.workspace when the dispatch and provider gates are enabled. See Execution workspace requests.

execution:
runtimeClassName: gvisor
nodeSelector:
sandbox-runtime: gvisor
tolerations:
- key: sandbox-runtime
operator: Equal
value: gvisor
effect: NoSchedule
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: sandbox-runtime
operator: In
values: ["gvisor"]
FieldTypeDescription
runtimeClassNamestringSelects a Kubernetes RuntimeClass such as gvisor or kata-qemu
nodeSelectormap[string]stringRestricts native worker Pods to nodes with matching labels
tolerationslistAllows native worker Pods onto tainted runtime-specific node pools
affinityobjectAdds Kubernetes affinity or anti-affinity rules for native worker Pods
workspaceobjectExecution-workspace provider request. With workspace dispatch enabled, provider: agent-sandbox (no templateRef) or provider: substrate (with a required infrastructure templateRef) hosts the Task's RuntimeSession in a workspace-provider-backed RuntimePool; everything else fails closed. Repository access always uses top-level Task.spec.workspace.

Resolution order:

  • Agent.spec.execution provides defaults for tasks that reference the Agent
  • Task.spec.execution overrides Agent defaults
  • runtimeClassName is a scalar override
  • nodeSelector, tolerations, and affinity replace Agent defaults when they are set on the Task

Execution workspace requests​

Task.spec.execution.workspace requests a physical execution-workspace provider for the Task's ACP RuntimeSession. With --acp-workspace-dispatch-enabled, provider: agent-sandbox (requires --agent-sandbox-enabled; templateRef must be omitted) or provider: substrate (requires --substrate-enabled; an infrastructure templateRef is required) executes the RuntimeSession in a dedicated workspace-provider-backed RuntimePool. Unsupported options (cleanupPolicy: retain, boot/pool/snapshot/hibernation, onDetach) fail closed before any workspace or RuntimePool demand, with the reason projected to Task.status.executionWorkspace. There is no worker Job fallback and no harness-v1 fallback. Top-level Task.spec.workspace remains the verified source/publication contract.

See Agent Sandbox Workspaces, Agent Substrate Workspaces, and ADRs 0024/0025 for the provider-neutral contract.

SubstrateActorPool​

SubstrateActorPool is an operator-owned pool of deterministic Substrate actors for pooled Task placement, MCP actor-backed Tools, and density reporting.

apiVersion: core.orka.ai/v1alpha1
kind: SubstrateActorPool
metadata:
name: codex-substrate-pool
spec:
templateRef:
name: orka-codex
namespace: ate-demo
workerPoolRef:
name: orka-workers
namespace: ate-demo
targetActors: 4
precreateActors: true
FieldTypeDefaultDescription
templateRef.namestringrequiredSubstrate ActorTemplate used for pool members.
templateRef.namespacestringPool namespaceNative Substrate Atespace containing the ActorTemplate.
workerPoolRef.namestringemptyOptional Substrate WorkerPool used for capacity and density reporting.
workerPoolRef.namespacestringPool namespaceNamespace containing the WorkerPool.
targetActorsinteger0Desired stateful actor count, capped at 1000. References from Tasks or Tools require at least 1.
precreateActorsbooleanfalsePre-create deterministic warm actors up to targetActors.

spec.templateRef is immutable. Orka records the first accepted native template UID in status.templateUID, including for empty pools. Replacing the native template under the same name requires a new pool.

For built-in OpenCode Agents, spec.model.name must use literal provider/model form and both spec.model.contextWindow and spec.model.maxTokens are required positive reviewed ceilings, with contextWindow > maxTokens. They are included in the immutable runtime profile; Orka does not discover or guess them from a mutable model catalog.

Provider fallback chain​

You can configure fallback providers that are automatically tried when the primary provider fails (e.g., due to auth errors, provider outages, or rate limiting). Fallbacks are configured on the Agent CRD's spec.model.fallbacks field.

apiVersion: core.orka.ai/v1alpha1
kind: Agent
metadata:
name: resilient-agent
spec:
providerRef:
name: my-openai
model:
name: gpt-4o
fallbacks:
- providerRef: my-anthropic
model: claude-sonnet-4-20250514
- providerRef: my-azure-openai
model: gpt-4o

How fallbacks work​

  1. The primary provider is tried first with automatic retries (exponential backoff on 429/5xx errors).
  2. If the primary provider fails with an auth error (401/403), network error, or exhausts all retries, the first fallback provider is tried.
  3. Each fallback provider also gets automatic retries.
  4. If all providers fail, the last error is returned.

Fallback fields​

FieldTypeRequiredDescription
providerRefstringYesName of a Provider CRD to fall back to
modelstringNoModel to use with this provider. If empty, uses the provider's defaultModel

Notes​

  • Fallbacks are only supported on Agent-based tasks. Agent-less tasks get retries only.
  • Each fallback provider must have its own Provider CRD with a valid secret reference.
  • Rate-limited providers (429 responses) are temporarily cooled down and skipped in subsequent requests.
  • Streaming requests are retried/failed over only on the initial connection — mid-stream failures are not retried.

Agent (with runtime)​

Agent configuration for the supported built-in ACP runtime profiles: Claude, Codex, Copilot, and OpenCode. Built-in ACP Agents do not reference provider Secrets; RuntimePools reach your configured model gateway through Orka's authenticated provider proxy.

apiVersion: core.orka.ai/v1alpha1
kind: Agent
metadata:
name: claude-agent
spec:
model:
name: "claude-sonnet-4-20250514"
systemPrompt:
inline: "You are a senior software engineer."
runtime:
type: claude # or "codex" / "copilot" / "opencode"
contractVersion: orka.harness.v2
defaultMaxTurns: 50
defaultAllowBash: true
defaultAllowedTools:
- Read
- Write
- Edit
- Bash

OpenCode Agents must omit spec.systemPrompt because the runtime cannot enforce Agent-level prompts; put instructions in each Task's spec.prompt instead. OpenCode model IDs use provider/model form, for example openai/gpt-5.4, and require reviewed contextWindow and maxTokens values.

Operator-owned runtimes outside the built-in set can use orka.harness.v2 AgentRuntime registration and conformance. A current-generation ready, strict-governed registration can be selected through runtime.runtimeRef; Orka freezes and revalidates its endpoint, profile, authentication authority, and observed runtime identity for dispatch.

Agent runtime tasks reference an Agent with runtime configured:

apiVersion: core.orka.ai/v1alpha1
kind: Task
metadata:
name: code-review
spec:
type: agent
agentRef:
name: claude-agent
prompt: "Review the code in this repo for security issues. Do not modify files."
workspace:
intent: read
gitRepo: "https://github.com/example/repo.git"
branch: main
# readCredentialRef:
# name: repository-read
# subPath: "services/api"
agentRuntime:
maxTurns: 100
allowBash: true
allowedTools:
- Read
- Write
- Edit
- Bash
- Glob
- Grep

Skill​

Reusable skill definitions (Agent Skills standard) that are referenced by Agents and AI Tasks.

apiVersion: core.orka.ai/v1alpha1
kind: Skill
metadata:
name: skill-researcher
labels:
orka.ai/category: "research"
spec:
displayName: "Research Methodology"
description: "Structured research workflow and source validation guidance"
version: "1.0.0"
author: "platform-team"
tags: ["research", "analysis"]
content:
inline: |
# Research Skill
Use primary sources and cite references.
files:
templates/checklist.md: |
- [ ] Validate source credibility
- [ ] Cross-check key claims
# source tracks where a skill was imported from (for updates)
# source:
# github: "anthropics/skills"
# skillName: "researcher"
status:
phase: Ready
contentHash: sha256:...

Tool​

Custom tool definitions for agents. Tools can call plain HTTP endpoints or MCP servers hosted in durable Substrate actors. Plain HTTP tools require http.url and support header-based or body-based auth injection.

This example uses a placeholder catalog API. Replace its URL and Secret reference with your service's values.

apiVersion: core.orka.ai/v1alpha1
kind: Tool
metadata:
name: catalog-search
spec:
description: "Search a product catalog"
parameters:
type: object
properties:
query:
type: string
description: "Search query"
required: ["query"]
http:
url: "https://catalog.example.com/search"
method: POST
timeout: 30s
authSecretRef:
name: catalog-api-key
key: api-key
authInject: body # "header" (Bearer token) or "body" (JSON key)
authBodyKey: api_key # JSON key name when authInject=body

MCP actor-backed tools can also set http.authSecretRef for transport auth. For these tools, http.url may be omitted because Orka uses the resolved actor endpoint from Tool status. MCP transport auth must use header injection; authInject: body is only valid for plain HTTP tools because MCP call arguments are forwarded to the MCP server as tool input.

Example MCP actor-backed Tool:

apiVersion: core.orka.ai/v1alpha1
kind: Tool
metadata:
name: repo-inspector
spec:
description: "Inspect repository metadata through an MCP server"
parameters:
type: object
properties:
message:
type: string
required:
- message
mcp:
path: /mcp
substrateActor:
templateRef:
name: orka-mcp
namespace: ate-demo
poolRef:
name: mcp-substrate-pool
boot: true

MCP actor-backed Tools require mcp.substrateActor.templateRef.name. mcp.path defaults to /mcp, poolRef is optional, and boot only affects the first actor resume. spec.http may be omitted unless the resolved actor endpoint needs transport auth settings; when spec.http is present for MCP auth only, omit http.url.

URL path interpolation​

Tool CRD URLs can contain {{paramName}} placeholders that are replaced with parameter values at runtime. Interpolated values are URL path-escaped, and the matching parameters are removed from the request body. This is useful for REST APIs that require path parameters.

apiVersion: core.orka.ai/v1alpha1
kind: Tool
metadata:
name: github-merge-pr
spec:
description: "Merge a GitHub pull request"
parameters:
type: object
properties:
owner:
type: string
repo:
type: string
pull_number:
type: integer
merge_method:
type: string
enum: [merge, squash, rebase]
required: [owner, repo, pull_number]
http:
url: "https://api.github.com/repos/{{owner}}/{{repo}}/pulls/{{pull_number}}/merge"
method: PUT
authSecretRef:
name: github-token
key: token
authInject: header

In this example, owner, repo, and pull_number are interpolated into the URL path and removed from the JSON body. Only merge_method is sent in the request body.

Provider​

LLM provider configuration with credentials. Supports Anthropic, OpenAI, and Azure OpenAI.

apiVersion: core.orka.ai/v1alpha1
kind: Provider
metadata:
name: anthropic-prod
spec:
type: anthropic # or "openai", "azure-openai"
secretRef:
name: anthropic-secret
key: api-key
baseURL: "" # optional custom endpoint for proxies
defaultModel: claude-sonnet-4-20250514
# Azure-specific (only for type: azure-openai)
# azure:
# deploymentName: my-deployment
# apiVersion: "2024-02-15-preview"

Helm chart​

Key configuration values for the Helm chart:

ParameterDefaultDescription
controller.replicas1Controller replicas
controller.image.repositoryghcr.io/orka-agents/orkaController image
controller.image.tagrelease versionController image tag.
controller.image.digest""Optional SHA256 digest; takes precedence over tag.
controller.modeharness-v2Static agent execution mode: harness-v1 or harness-v2. Select v1 explicitly for a compatibility release; a release never serves both or changes mode in place.
controller.watchNamespaceHelm release namespaceMust match the release namespace, labeled orka.ai/controller-mode with the matching mode.
controller.agentExecutionSnapshot.existingSecret""Your own snapshot encryption Secret. Empty renders an empty <release>-agent-execution-snapshot that the controller fills on first start. Immutable after install. See Snapshot encryption key.
controller.agentExecutionSnapshot.keykeyItem inside that Secret holding 32 raw bytes or their base64 encoding. Immutable after install.
controller.enforceNamespaceIsolationtrueRestrict namespace-bound API callers and default Helm RBAC to their namespace
service.port8080Controller Service port used by controller and Publisher in-cluster URLs.
controller.apiPort8080Controller container listener and Service target port.
controller.metricsPort8081Metrics endpoint port
controller.healthPort8082Health probe port
controller.logLevelinfoLog level (debug/info/warn/error)
controller.acpRuntime.namespaceorka-runtimesNamespace for controller-owned RuntimePool workloads.
controller.acpRuntime.providerProxyNamespace""Compatibility guard for the chart-managed provider proxy. Leave empty or set exactly to the Helm release namespace; any other nonempty value is rejected when the proxy is enabled.
controller.acpRuntime.codexImagerelease image tagCodex ACP image tag or digest; an empty string disables it.
controller.acpRuntime.claudeImagerelease image tagClaude ACP image tag or digest; an empty string disables it.
controller.acpRuntime.copilotImagerelease image tagGitHub Copilot ACP image tag or digest; an empty string disables it.
controller.acpRuntime.opencodeImagerelease image tagOpenCode ACP image tag or digest; an empty string disables it.
controller.acpRuntime.upgradeDrain.*enabledTwo-phase planned-upgrade admission closure and RuntimePool drain settings.
harnessV1.image.digest""Required immutable wrapper image digest for a harness-v1 release.
harnessV1.auth.existingSecret""Dedicated v1 wrapper bearer Secret, separate from its TLS Secret. Never share it with v2.
harnessV1.tls.existingSecret""Dedicated v1 wrapper Secret containing tls.crt, tls.key, and ca.crt.
harnessV1.tls.rolloutNonce""Non-secret revision marker for certificate renewal without changing the TLS Secret name.
providerProxy.enabledfalseEnable the authenticated proxy when connecting built-in coding agents to your model gateway. Installation and AI-worker tasks do not require it.
providerProxy.upstreamBaseURL""Your gateway's HTTP(S) endpoint. Required when the proxy is enabled. Credentials, queries, and fragments are forbidden in the URL.
providerProxy.egress[]NetworkPolicy egress rules allowing the proxy to reach your gateway. Leave empty when upstreamBaseURL names an in-cluster Service; the chart then derives the rule from that Service at install time. Required for external gateways, named target ports, or offline renders. DNS access is always added.
providerProxy.auth.existingSecret""Existing current/optional-overlap proxy bearer Secret. RuntimePool copies are controller-managed.
providerProxy.tokenReloadInterval5sAtomic projected-Secret reload interval. Invalid generations fail readiness and forwarding closed.
publisher.enabledtrueDeploy the separate clean-room Workspace/Publisher service.
publisher.image.repository / publisher.image.tagworkspace publisher image / release versionPublisher image. Set publisher.image.digest to override the tag.
publisher.allowedSCMHostsgithub.comExact lower-case SCM hosts accepted by both Publisher validation and the SCM egress proxy.
publisher.auth.existingSecret""Existing controller-auth/capability Secret for publisher operations.
publisher.auth.rolloutNonce""Non-secret revision marker that restarts controller and Publisher during coordinated publisher-auth Secret rotation.
scmEgressProxy.enabledtrueRequire all Publisher Git and forge HTTPS traffic to traverse the dedicated authenticated proxy. The chart rejects publisher.enabled=true when this is false.
scmEgressProxy.auth.existingSecret""Existing Secret shared only by Publisher and proxy. Its token must contain 32-256 RFC 3986 unreserved characters.
scmEgressProxy.auth.rolloutNonce""Non-secret revision marker that restarts Publisher and SCM proxy during coordinated proxy-auth Secret rotation.
scmEgressProxy.maxTunnelBytes1073741824Maximum bytes allowed in each CONNECT tunnel direction.
scmEgressProxy.maxConcurrent8Maximum concurrent forward requests and CONNECT tunnels.
webhooks.tls.existingSecret""Your own TLS Secret for the admission webhooks. Empty lets the controller issue and renew a self-signed certificate in <release>-webhook-tls. See Webhook certificate.
webhooks.tls.certKey / webhooks.tls.privateKeyKeytls.crt / tls.keyCertificate and private-key keys inside the webhook TLS Secret.
webhooks.caBundle""Base64-encoded PEM CA bundle for an operator-supplied certificate. Leave empty with controller-issued TLS, which injects its own CA, or when webhooks.caInjectionAnnotations configures an injector.
webhooks.caInjectionAnnotations{}CA-injection annotations (for example cert-manager) placed on the chart ValidatingWebhookConfiguration. With webhooks.tls.existingSecret set, rendering fails unless this or webhooks.caBundle is set.
webhooks.timeoutSeconds10Admission webhook timeout.
controller.agentSandbox.enabledfalseEnable experimental workspace-backed execution for agent Tasks that set execution.workspace
controller.agentSandbox.routerUrl""Optional upstream agent-sandbox router base URL used for workspace claims
controller.agentSandbox.defaultTemplate""Default agent-sandbox SandboxWarmPool name when a Task omits templateRef.name
controller.agentSandbox.warmPoolPolicydisabledLegacy compatibility setting: disabled or template; v1 claims use SandboxWarmPool references
controller.agentSandbox.namespaceStrategytaskSandbox resource namespace strategy: task or controller
controller.agentSandbox.claimTimeout2mTimeout for workspace claim and readiness operations
controller.agentSandbox.commandTimeout30mTimeout for agent runtime execution inside the sandbox
controller.agentSandbox.cleanupPolicydeleteLegacy setting, ignored by harness-v2 ACP. An omitted Task cleanupPolicy defaults to delete; retain is rejected.
workers.ai.image.repositoryghcr.io/orka-agents/orka/ai-workerAI worker image
workers.general.image.repositoryghcr.io/orka-agents/orka/general-workerGeneral worker image
service.typeClusterIPService type
client.createtrueCreate client ServiceAccount for API access
client.nameorka-clientClient ServiceAccount name
client.namespaceHelm release namespaceClient ServiceAccount namespace; must match controller.watchNamespace.

Image overrides​

Released charts include matching version tags for the controller, workers, Publisher, and coding-agent runtimes. For controller, worker, and Publisher images, set image.tag to choose another version or image.digest to pin a SHA256 digest. A digest takes precedence over the tag.

Runtime fields such as controller.acpRuntime.codexImage accept a full image reference with a tag or digest. The controller resolves tags to digests once at startup, so running sessions use fixed images. If any resolution fails, the controller logs the error, disables every coding-agent runtime for that process, and keeps running: AI and container Tasks work, agent Tasks fail closed as unavailable. Fix registry access or pin digests, then restart the controller. An explicit digest skips this lookup.

Tag resolution requires controller HTTPS access to a registry that allows anonymous pulls. Use digest references for private registries or installations without registry access from the controller. Cluster nodes still need access to pull the configured images. Set an individual runtime image to an empty string to disable that provider.

Snapshot encryption key​

The controller encrypts agent execution snapshots at rest with an AES-256 key mounted from a Secret. By default the chart renders an empty Secret named <release>-agent-execution-snapshot and the controller mints 32 random bytes into it on its first start, then reuses them on every later start. The chart never renders key material, so Helm release history and dry-run output contain no secrets, and Helm leaves the controller-written data alone on upgrade. The Secret is kept on helm uninstall so a restored data volume can still be decrypted. Back it up with the data volume.

If the controller finds its SQLite store already present but the Secret empty, the key was lost. It refuses to start rather than mint a new key that could not read the existing records. Restore the Secret from backup, or delete the data volume to start over.

To bring your own key, create the Secret before installing and set controller.agentExecutionSnapshot.existingSecret. The item must hold exactly 32 raw bytes or their base64 encoding:

kubectl -n orka-system create secret generic orka-agent-snapshot-key \
--from-literal=key="$(openssl rand -base64 32)"
helm install orka orka/orka --namespace orka-system \
--set-string controller.agentExecutionSnapshot.existingSecret=orka-agent-snapshot-key \
--set-string controller.agentExecutionSnapshot.key=key \
...

The Secret name, item key, and key material are immutable for the life of the release. Helm rejects an upgrade that changes the name or item, including a switch between the generated Secret and your own. It cannot detect changed bytes under the same name, so never rotate the material in place, and never run helm upgrade --force, which replaces the Secret with the chart's empty version.

Webhook certificate​

Orka validates its resources through fail-closed Kubernetes admission webhooks served by the controller. Those webhooks need a serving certificate that the API server trusts. The chart supports two ways to provide it.

Controller-managed, the default. With webhooks.tls.existingSecret empty, the chart renders an empty Secret named <release>-webhook-tls and the controller fills it using cert-controller, the library Gatekeeper uses for the same job. It mints a ten-year self-signed CA and a one-year serving certificate, renews the serving certificate before expiry, and writes the CA into the caBundle of the release's ValidatingWebhookConfiguration. The Secret is the source of truth and is kept on helm uninstall, so a reinstall under the same release name reuses the CA. The chart never renders certificate material. The Secret is mounted into the controller Pod, and the Pod reports ready only after the kubelet has projected the certificate and the CA is injected, which can take about a minute on a fresh install. Never run helm upgrade --force, which replaces the Secret with the chart's empty version. This grants the controller list and watch on all ValidatingWebhookConfigurations, which cannot be name-scoped, plus update on its own. Because that update is whole-object, a ValidatingAdmissionPolicy installed with the release lets the controller's ServiceAccount change nothing on that configuration except each webhook's caBundle, so a compromised controller cannot repoint or widen its own admission webhooks.

Operator-supplied. Set webhooks.tls.existingSecret to a Secret holding tls.crt and tls.key valid for <release>-webhook.<namespace>.svc, and supply the CA either as webhooks.caBundle or through webhooks.caInjectionAnnotations. The controller then only reads the mounted files and never touches the webhook configuration.

webhooks:
tls:
existingSecret: orka-webhook-tls
caInjectionAnnotations:
cert-manager.io/inject-ca-from-secret: orka-system/orka-webhook-tls

For a one-off self-signed certificate without cert-manager:

(
umask 077
openssl req -x509 -newkey rsa:2048 -nodes -sha256 -days 365 \
-keyout tls.key -out tls.crt \
-subj '/CN=orka-webhook.orka-system.svc' \
-addext 'subjectAltName=DNS:orka-webhook.orka-system.svc,DNS:orka-webhook.orka-system.svc.cluster.local' &&
cp tls.crt ca.crt
)
kubectl -n orka-system create secret generic orka-webhook-tls \
--type=kubernetes.io/tls \
--from-file=tls.crt=tls.crt --from-file=tls.key=tls.key --from-file=ca.crt=ca.crt
rm -f tls.key tls.crt ca.crt
helm install orka orka/orka --namespace orka-system \
--set-string webhooks.tls.existingSecret=orka-webhook-tls \
--set-string webhooks.caBundle="$(kubectl -n orka-system get secret orka-webhook-tls -o jsonpath='{.data.ca\.crt}')" \
...

Switching between the two modes is an ordinary helm upgrade. A certificate for any other name is rejected by the API server at admission time, not at install time.

Helm authentication Secret rotation​

The Publisher, SCM proxy, and controller clients load these credentials at process startup. Update the Secret and bump its corresponding nonce in the same Helm upgrade: use publisher.auth.rolloutNonce for the publisher-auth Secret, and scmEgressProxy.auth.rolloutNonce for the SCM proxy-auth Secret. The publisher nonce is applied only to controller and Publisher Pod templates; the SCM nonce is applied only to Publisher and SCM proxy Pod templates. Nonces are safe revision strings, not Secret values. Coordinated rotation can briefly fail closed while Pods roll but prevents stale or split credential generations from persisting.

Harness v1 certificate rotation​

The harness-v1 wrapper keeps bearer and TLS credentials in separate Secrets. harnessV1.auth.existingSecret contains only the bearer token and is immutable while the wrapper Deployment exists. harnessV1.tls.existingSecret contains tls.crt, tls.key, and ca.crt.

Changing the TLS Secret name triggers a drained wrapper restart. For certificate renewal under the same name, update the TLS Secret and bump harnessV1.tls.rolloutNonce in the same Helm upgrade. The hook drains the live wrapper before restarting the wrapper and controller.

Keep the updated ca.crt able to verify the certificate served during the drain. Alternatively, use a new TLS Secret name so the hook can mount the prior CA.

Canonical Kustomize overlay​

Direct Kustomize deployments must use:

kubectl apply -k config/acp-production

config/acp-production composes the CRD-free config/acp-workload base with the cross-namespace Vekil ingress NetworkPolicy. Applying config/default alone omits the boundary that prevents runtime Pods from bypassing the authenticated provider proxy. Configure the required system Secrets and digest-pinned controller, Publisher, Codex, Claude, Copilot, and OpenCode images before applying the overlay. make deploy validates those image references and applies the equivalent resource set.

ACP storage split​

The --store-backend=sqlite flag does not make SQLite authoritative for ACP control transitions. ControllerEpoch, PromptAttempt, RuntimeSessionControl, BranchClaim, Publication, and ExternalEffect status plus coordination Leases are the control authority. SQLite stores transcript/SessionTurn payloads, deferred outbox projections, and artifacts behind those Kubernetes fences.

For the Kustomize deployment, create the proxy-auth Secret before applying config/acp-production (the Secret is intentionally not stored in Git):

token="$(openssl rand -hex 32)"
kubectl -n orka-system create secret generic scm-egress-proxy-auth \
--from-literal=token="$token"
unset token

The token is an ingress credential for the proxy only; it is not a Git or forge credential. config/publisher sets HTTPS_PROXY/NO_PROXY, and the Publisher copies only those validated proxy variables into its otherwise empty Git subprocess environments. The Publisher NetworkPolicy has no public 0.0.0.0/0 rule. Only the SCM egress proxy may reach public port 443.

cmd/orka-workspace-publisher also requires ORKA_PUBLISHER_ARTIFACT_AUTHORIZATION_BROKER_URL, ORKA_PUBLISHER_CREDENTIAL_BROKER_URL, and ORKA_PUBLISHER_SCM_EGRESS_PROXY_REQUIRED=true in normal startup.

ORKA_PUBLISHER_ALLOW_DEVELOPMENT_FALLBACKS is for isolated local tests only

It turns on the legacy local artifact signing key, a filesystem credential root, and proxy-less mode — three of the boundaries that keep publication credentials away from the agent. Never set it in a cluster manifest.

When ORKA_PUBLISHER_PUBLISH_TIMEOUT is raised above the default on the Publisher, set the same value on the controller Deployment as well. The controller bounds each publisher-backed external-effect call and sizes the effect's ledger lease from this timeout plus a settlement margin; without the controller-side value, publisher operations are clamped to the default four-minute call bound. Brokered custom-Tool calls always keep the fixed four-minute clamp that their Tool descriptors were admitted under.

Helm CRD lifecycle​

CRD behavior is not controlled through chart values. A fresh install creates all CRDs in the chart unless --skip-crds is used. Because CRDs are cluster-scoped, designate one lifecycle owner and use --skip-crds for other Orka releases. Helm does not update CRDs during helm upgrade; apply the CRDs from the exact target chart before upgrading the controller. Helm retains CRDs and Orka custom resources on uninstall. See the Helm CRD lifecycle guide.

Harness v1 and v2 use separate releases, endpoints, watched namespaces, RBAC, Leases, stores, and data planes. They do not migrate Tasks or continue Sessions across modes. See Operating harness v1 and v2 on one cluster.

Context-token flags can also be configured through Helm under controller.contextToken. For example:

controller:
contextToken:
profile: transaction-token
issuer: https://issuer.example.com
audience: orka
headers: Txn-Token
authzMode: enforce
scopes:
taskCreate: orka:tasks:create
providerUse: orka:providers:use
toolUse: orka:tools:use
secretCredentialRead: orka:secrets:credentials:read
monitorRead: orka:monitors:read
monitorWrite: orka:monitors:write
monitorOperate: orka:monitors:operate
connectorRead: orka:connectors:read
connectorManage: orka:connectors:manage
gatewayRead: orka:gateways:read
gatewayOperate: orka:gateways:operate
tts:
endpoint: https://tts.example.com/oauth/token
audience: orka-workers
timeout: 5s
tokenSource: serviceAccount
childScope: orka:tasks:create
outboundScope: orka:tools:use
childTokenTTL: 5m
toolTokenTTL: 2m
outboundAccess:
trustedGatewayServices:
- gateway-system/agentgateway:8080
trustedTokenEndpointServices:
- identity-system/token-service:8443

The Helm keys mirror the controller flags: for example, controller.contextToken.jwksUrl renders --context-token-jwks-url, controller.contextToken.scopes.secretRead renders --context-token-secret-read-scopes, controller.contextToken.scopes.secretCredentialRead renders --context-token-secret-credential-read-scopes, controller.contextToken.scopes.monitorRead renders --context-token-monitor-read-scopes, controller.contextToken.scopes.gatewayRead renders --context-token-gateway-read-scopes, and controller.contextToken.tts.toolTokenTTL renders --context-token-tool-token-ttl.

See charts/orka/values.yaml for the full list.

Controller flags​

FlagDefaultDescription
--api-port8080REST API server port
--gateway-enabledtrueEnable generic gateway reconciliation and ingress
--connectors-enabledfalseEnable per-user connector reconciliation (ConnectorProvider and Connection). Requires Task provenance admission (--task-provenance-admission-enabled or --task-provenance-admission-external); the controller refuses to start otherwise, because connector use trusts spec.requestedBy only when the API server provably stamped it. Env: ORKA_CONNECTORS_ENABLED
--connectors-allow-private-endpointsemptyLocal fixtures only. Accept private, loopback, and cluster-local connector provider endpoints and, for Tools behind a connection-mode outbound access policy, private tool endpoints (HTTPS is still required), and let linked-account requests reach them over the same hardened transport (no proxy, verified TLS). A person's token would be sent to such an address, so the flag takes only the literal i-understand-tokens-may-leave-the-cluster, and the controller refuses to start with it unless --connector-callback-base-url is a plain-http localhost origin, which no production deployment has. The Live Connectors E2E uses it for its in-cluster fake provider. Env: ORKA_CONNECTORS_ALLOW_PRIVATE_ENDPOINTS
--connector-callback-base-urlORKA_CONNECTOR_CALLBACK_BASE_URL env or ""Required with --connectors-enabled. Absolute https origin (scheme and host only, no path) the OAuth provider redirects back to; the provider must register exactly this origin plus /api/v1/connections/callback. Plain http is accepted only for localhost. Helm: controller.connectors.callbackBaseUrl. Sealed linked-account credentials live in the controller store, which the chart already requires to be persistent in every mode.
--gateway-pending-per-session100Maximum pending gateway events per Session
--gateway-interim-messages-per-task10Lifetime cap for distinct accepted interim messages per Task; failed/expired messages count, retries do not. Helm: controller.gateway.interimMessagesPerTask
--gateway-max-records-per-gateway1000Maximum retained accepted/dead-letter event records per Gateway before ingress is throttled
--gateway-max-rejected-records-per-gateway250Separate audit budget for rejected events so unauthorized traffic cannot consume operational capacity
--gateway-event-expiry24hQueue and delivery retry expiry
--gateway-terminal-retention720hTerminal event and delivery retention
--gateway-delivery-timeout15sOne synchronous adapter delivery timeout
--gateway-delivery-max-attempts10Delivery attempts before dead-lettering
--gateway-claim-lease1mEvent and delivery claim lease
--gateway-poll-interval500msDispatcher and delivery poll interval
--gateway-batch-size25Maximum gateway records processed per iteration
--controller-mode / ORKA_CONTROLLER_MODErequiredStatic controller mode: harness-v1 or harness-v2. dual, auto, and drain modes are rejected.
--watch-namespacerequiredOne non-empty watched namespace. Its orka.ai/controller-mode label must match; the controller adds it when absent.
--claim-namespace-modetrueLabel an unlabeled watched namespace with the controller's mode on first start. Set false to require an operator-applied label.
--enforce-namespace-isolationfalseRestrict users to their ServiceAccount's namespace
--max-tasks-per-namespace0Max active tasks per namespace (0 = unlimited)
--agent-sandbox-enabledORKA_AGENT_SANDBOX_ENABLED env or falseAdmit the agent-sandbox execution-workspace provider for agent Tasks that set execution.workspace
--acp-workspace-dispatch-enabledORKA_ACP_WORKSPACE_DISPATCH_ENABLED env or falseAdmit workspace-provider-backed ACP RuntimeSession dispatch; when false, workspace-backed agent Tasks fail closed
--agent-sandbox-router-urlORKA_AGENT_SANDBOX_ROUTER_URL env or ""Optional upstream agent-sandbox router base URL used for workspace claims
--agent-sandbox-default-templateORKA_AGENT_SANDBOX_DEFAULT_TEMPLATE env or ""Default agent-sandbox SandboxWarmPool name when a Task omits templateRef.name
--agent-sandbox-warm-pool-policyORKA_AGENT_SANDBOX_WARM_POOL_POLICY env or disabledLegacy compatibility setting: disabled or template; v1 claims use SandboxWarmPool references
--agent-sandbox-namespace-strategyORKA_AGENT_SANDBOX_NAMESPACE_STRATEGY env or taskSandbox resource namespace strategy: task or controller
--agent-sandbox-claim-timeoutORKA_AGENT_SANDBOX_CLAIM_TIMEOUT env or 2mTimeout for workspace claim and readiness operations
--agent-sandbox-command-timeoutORKA_AGENT_SANDBOX_COMMAND_TIMEOUT env or 30mTimeout for agent runtime execution inside the sandbox
--agent-sandbox-cleanup-policyORKA_AGENT_SANDBOX_CLEANUP_POLICY env or deleteLegacy setting, ignored by harness-v2 ACP. An omitted Task cleanupPolicy defaults to delete; retain is rejected.
--controller-url""Base URL workers use to reach the controller API (e.g., http://orka-api.orka-system.svc:8080). Required for worker result callbacks and session transcript fetching
--oidc-issuerORKA_OIDC_ISSUER env or ""OIDC issuer URL for external API bearer token validation. Requires --oidc-audience when set
--oidc-audienceORKA_OIDC_AUDIENCE env or ""Expected OIDC audience for external API bearer tokens. Requires --oidc-issuer when set
--oidc-jwks-urlORKA_OIDC_JWKS_URL env or ""Optional JWKS URL. When empty, Orka discovers it from the issuer metadata
--oidc-allowed-subjectsORKA_OIDC_ALLOWED_SUBJECTS env or ""Required comma-separated OIDC subject allowlist patterns when OIDC is enabled
--oidc-namespaceORKA_OIDC_NAMESPACE env or defaultNamespace assigned to authorized OIDC callers for namespace isolation
--context-token-profileORKA_CONTEXT_TOKEN_PROFILE env or ""Context-token profile for external API requests. Currently supports transaction-token
--context-token-issuerORKA_CONTEXT_TOKEN_ISSUER env or ""Context-token issuer URL. Requires --context-token-profile and --context-token-audience when set
--context-token-audienceORKA_CONTEXT_TOKEN_AUDIENCE env or ""Expected context-token audience. Requires --context-token-profile and --context-token-issuer when set
--context-token-jwks-urlORKA_CONTEXT_TOKEN_JWKS_URL env or ""Optional context-token JWKS URL. For transaction-token, defaults to <issuer>/.well-known/jwks.json
--context-token-headersORKA_CONTEXT_TOKEN_HEADERS env or ""Comma-separated context-token header locations. Use Header for raw tokens or Header:Scheme for scheme-prefixed tokens. The transaction-token default is Txn-Token
--context-token-authz-modeORKA_CONTEXT_TOKEN_AUTHZ_MODE env or ""Context-token authorization mode: off, audit, or enforce. Empty defaults to off
--context-token-task-create-scopesORKA_CONTEXT_TOKEN_TASK_CREATE_SCOPES env or ""Comma-separated scopes authorizing Task creation. Defaults to orka:tasks:create
--context-token-task-read-scopesORKA_CONTEXT_TOKEN_TASK_READ_SCOPES env or ""Comma-separated scopes authorizing Task reads and related data. Defaults to orka:tasks:get
--context-token-task-list-scopesORKA_CONTEXT_TOKEN_TASK_LIST_SCOPES env or ""Comma-separated scopes authorizing Task listing. Defaults to orka:tasks:list
--context-token-task-delete-scopesORKA_CONTEXT_TOKEN_TASK_DELETE_SCOPES env or ""Comma-separated scopes authorizing Task deletion. Defaults to orka:tasks:delete
--context-token-tool-read-scopesORKA_CONTEXT_TOKEN_TOOL_READ_SCOPES env or ""Comma-separated scopes authorizing Tool reads. Defaults to orka:tools:read
--context-token-tool-use-scopesORKA_CONTEXT_TOKEN_TOOL_USE_SCOPES env or ""Comma-separated scopes authorizing Orka-managed chat/OpenAI/Anthropic tool execution. Defaults to orka:tools:use
--context-token-provider-use-scopesORKA_CONTEXT_TOKEN_PROVIDER_USE_SCOPES env or ""Comma-separated scopes authorizing chat/OpenAI/Anthropic model-provider use and model listing. Defaults to orka:providers:use
--context-token-secret-read-scopesORKA_CONTEXT_TOKEN_SECRET_READ_SCOPES env or ""Comma-separated scopes authorizing Secret metadata reads. Defaults to orka:secrets:read
--context-token-secret-credential-read-scopesORKA_CONTEXT_TOKEN_SECRET_CREDENTIAL_READ_SCOPES env or ""Comma-separated scopes authorizing Secret data or ServiceAccount tokens as outbound credentials. Defaults to orka:secrets:credentials:read
--context-token-agent-read-scopesORKA_CONTEXT_TOKEN_AGENT_READ_SCOPES env or ""Comma-separated scopes authorizing Agent reads. Defaults to orka:agents:read
--context-token-agent-write-scopesORKA_CONTEXT_TOKEN_AGENT_WRITE_SCOPES env or ""Comma-separated scopes authorizing Agent writes. Defaults to orka:agents:write
--context-token-memory-read-scopesORKA_CONTEXT_TOKEN_MEMORY_READ_SCOPES env or ""Comma-separated scopes authorizing memory reads. Defaults to orka:memory:read
--context-token-memory-write-scopesORKA_CONTEXT_TOKEN_MEMORY_WRITE_SCOPES env or ""Comma-separated scopes authorizing memory writes. Defaults to orka:memory:write
--context-token-session-read-scopesORKA_CONTEXT_TOKEN_SESSION_READ_SCOPES env or ""Comma-separated scopes authorizing session reads. Defaults to orka:sessions:read
--context-token-session-write-scopesORKA_CONTEXT_TOKEN_SESSION_WRITE_SCOPES env or ""Comma-separated scopes authorizing session writes/deletes. Defaults to orka:sessions:write
--context-token-security-read-scopesORKA_CONTEXT_TOKEN_SECURITY_READ_SCOPES env or ""Comma-separated scopes authorizing security scan reads. Defaults to orka:security:read
--context-token-security-write-scopesORKA_CONTEXT_TOKEN_SECURITY_WRITE_SCOPES env or ""Comma-separated scopes authorizing security scan creates, updates, deletes, and other mutations. Defaults to orka:security:write
--context-token-monitor-read-scopesORKA_CONTEXT_TOKEN_MONITOR_READ_SCOPES env or ""Comma-separated scopes authorizing repository monitor reads. Defaults to orka:monitors:read
--context-token-monitor-write-scopesORKA_CONTEXT_TOKEN_MONITOR_WRITE_SCOPES env or ""Comma-separated scopes authorizing repository monitor create, update, and delete operations. Defaults to orka:monitors:write
--context-token-monitor-operate-scopesORKA_CONTEXT_TOKEN_MONITOR_OPERATE_SCOPES env or ""Comma-separated scopes authorizing repository monitor manual runs. Defaults to orka:monitors:operate
--context-token-connector-read-scopesORKA_CONTEXT_TOKEN_CONNECTOR_READ_SCOPES env or ""Comma-separated scopes authorizing a person to read their own connector Connections. Defaults to orka:connectors:read
--context-token-connector-manage-scopesORKA_CONTEXT_TOKEN_CONNECTOR_MANAGE_SCOPES env or ""Comma-separated scopes authorizing a person to link, update, and disconnect their own connector Connections. Defaults to orka:connectors:manage
--context-token-skill-read-scopesORKA_CONTEXT_TOKEN_SKILL_READ_SCOPES env or ""Comma-separated scopes authorizing Skill reads. Defaults to orka:skills:read
--context-token-skill-write-scopesORKA_CONTEXT_TOKEN_SKILL_WRITE_SCOPES env or ""Comma-separated scopes authorizing Skill writes. Defaults to orka:skills:write
--context-token-gateway-read-scopesORKA_CONTEXT_TOKEN_GATEWAY_READ_SCOPES env or ""Comma-separated scopes authorizing gateway resource and ledger reads. Defaults to orka:gateways:read
--context-token-gateway-operate-scopesORKA_CONTEXT_TOKEN_GATEWAY_OPERATE_SCOPES env or ""Comma-separated scopes authorizing dead-lettered delivery retries. Defaults to orka:gateways:operate
--context-token-tts-endpointORKA_CONTEXT_TOKEN_TTS_ENDPOINT env or ""Exact transaction-token TTS OAuth endpoint for optional exchange/replacement
--context-token-tts-audienceORKA_CONTEXT_TOKEN_TTS_AUDIENCE env or ""Audience requested from transaction-token TTS exchanges
--context-token-tts-timeoutORKA_CONTEXT_TOKEN_TTS_TIMEOUT env or ""Timeout for transaction-token TTS exchanges. Defaults to 5s when TTS is enabled
--context-token-tts-token-sourceORKA_CONTEXT_TOKEN_TTS_TOKEN_SOURCE env or ""Subject token source for TTS exchanges: serviceAccount, incoming, or none. Defaults to serviceAccount when TTS is enabled
--context-token-subject-token-typeORKA_CONTEXT_TOKEN_SUBJECT_TOKEN_TYPE env or ""Subject token type for worker-side TTS exchanges. Workers default to TxToken subject tokens when empty
--context-token-child-scopeORKA_CONTEXT_TOKEN_CHILD_SCOPE env or ""Scope workers request for child delegated TxTokens when TTS is configured
--context-token-outbound-scopeORKA_CONTEXT_TOKEN_OUTBOUND_SCOPE env or ""Scope workers request for outbound HTTP Tool TxTokens when TTS is configured
--context-token-child-token-ttlORKA_CONTEXT_TOKEN_CHILD_TOKEN_TTL env or ""Requested TTL for child delegation TxTokens. Defaults to 5m when TTS is enabled
--context-token-tool-token-ttlORKA_CONTEXT_TOKEN_TOOL_TOKEN_TTL env or ""Requested TTL for outbound tool TxTokens. Defaults to 2m when TTS is enabled
--outbound-access-trusted-gateway-servicesORKA_OUTBOUND_ACCESS_TRUSTED_GATEWAY_SERVICES env or ""Comma-separated exact namespace/name:port cross-namespace gateway Service refs; wildcards are rejected
--outbound-access-trusted-token-endpoint-servicesORKA_OUTBOUND_ACCESS_TRUSTED_TOKEN_ENDPOINT_SERVICES env or ""Comma-separated exact namespace/name:port cross-namespace token endpoint Service refs; wildcards are rejected
--task-provenance-admission-enabledORKA_TASK_PROVENANCE_ADMISSION_ENABLED env or falseEnable validating admission that rejects untrusted direct Kubernetes Task writes to Orka-managed provenance fields (spec.requestedBy, spec.transaction, and transaction metadata labels/annotations)
--task-provenance-admission-externalORKA_TASK_PROVENANCE_ADMISSION_EXTERNAL env or falseDeclare that a separately deployed fail-closed Task provenance webhook protects coordination ancestry and workspace settlement metadata without hosting a manager webhook
--task-provenance-admission-trusted-usersORKA_TASK_PROVENANCE_ADMISSION_TRUSTED_USERS env or controller ServiceAccount usernamesComma-separated Kubernetes usernames trusted to set Orka-managed Task provenance fields
--task-provenance-admission-trusted-service-accountsORKA_TASK_PROVENANCE_ADMISSION_TRUSTED_SERVICE_ACCOUNTS env or configured AI/vendor worker ServiceAccountsComma-separated ServiceAccount names trusted in the target Task namespace to set Orka-managed Task provenance fields for child Task creation. Explicit values override the worker ServiceAccount defaults.
--ai-worker-imageghcr.io/orka-agents/orka/ai-worker:latestNative AI worker container image
--acp-runtime-namespace / ORKA_ACP_RUNTIME_NAMESPACEorka-runtimesNamespace for managed runtime Deployments, Services, Secrets, and policies.
--acp-provider-proxy-namespace / ORKA_ACP_PROVIDER_PROXY_NAMESPACE""Approved provider-proxy namespace selector. The Helm chart uses its release namespace when the proxy is enabled.
--acp-provider-proxy-base-url / ORKA_ACP_PROVIDER_PROXY_BASE_URLunsetAuthenticated provider-proxy URL injected into built-in RuntimePools.
--acp-provider-proxy-pod-labels / ORKA_ACP_PROVIDER_PROXY_POD_LABELSorka.ai/network-role=provider-auth-proxyExact Pod labels selected by RuntimePool egress policy.
--acp-provider-proxy-token-file / ORKA_ACP_PROVIDER_PROXY_TOKEN_FILEunsetController-mounted bearer file copied into generation-scoped immutable RuntimePool Secrets.
--acp-codex-runtime-image / ORKA_ACP_CODEX_RUNTIME_IMAGEunsetCodex runtime image with an explicit tag or SHA256 digest. Tags are resolved at startup.
--acp-claude-runtime-image / ORKA_ACP_CLAUDE_RUNTIME_IMAGEunsetClaude runtime image with an explicit tag or SHA256 digest. Tags are resolved at startup.
--acp-copilot-runtime-image / ORKA_ACP_COPILOT_RUNTIME_IMAGEunsetGitHub Copilot runtime image with an explicit tag or SHA256 digest. Tags are resolved at startup.
--acp-opencode-runtime-image / ORKA_ACP_OPENCODE_RUNTIME_IMAGEunsetOpenCode runtime image with an explicit tag or SHA256 digest. Tags are resolved at startup.
--general-worker-imageghcr.io/orka-agents/orka/general-worker:latestGeneral worker container image
--store-backendsqlitePayload/read-model backend. ACP control authority remains Kubernetes CRDs and Leases.
--store-path/data/orka.dbPath to the SQLite transcript/outbox/artifact database file.
--ai-worker-service-account-nameorka-ai-workerServiceAccount name for AI worker Jobs and dynamically ensured worker RBAC
--vendor-worker-service-account-nameorka-vendor-workerServiceAccount name for vendor/agent worker Jobs and dynamically ensured worker RBAC
--container-worker-service-account-nameorka-container-workerServiceAccount name for container worker Jobs and dynamically ensured worker RBAC
--chat-enabledtrueEnable the chat endpoint
--chat-provider""Default Provider CRD name for chat
--chat-model""Default model for chat
--chat-max-iterations50Max tool execution loops per chat request
--chat-max-duration30mMax wall-clock time per chat request
--chat-tool-timeout60sMax time for single tool execution
--chat-max-concurrent10Max concurrent chat sessions
--chat-max-tasks-per-turn5Max tasks created per chat turn
--chat-max-session-size512000Soft limit for session size before truncation (bytes)
--leader-electfalseEnable leader election. Static controller installations require true; the Lease is stored in the watched namespace.
--metrics-bind-address0Metrics endpoint address
--health-probe-bind-address:8081Health probe address
--metrics-securetrueServe metrics via HTTPS
--enable-http2falseEnable HTTP/2 for metrics and webhook servers
--enable-telemetry / --enable-tracingfalseEnable OpenTelemetry traces and metrics (requires worker-reachable OTLP endpoint for worker telemetry)

Workspace providers (workspace.orka.ai/v1alpha1)​

A workspace provider gives an agent a real machine to work on — a sandboxed container with a filesystem, rather than a plain Kubernetes Pod. Orka does not ship one; it talks to an external one. Two backends are supported: agent-sandbox and Substrate.

The API is installed alongside everything else but its controllers start off. Nothing below happens until you turn them on.

Turning it on​

FlagEnvironment variablePurpose
--enable-workspace-provider-apiORKA_ENABLE_WORKSPACE_PROVIDER_APIEnables the provider, class, pool, and workspace reconcilers.
--task-provenance-admission-enabled=true or --task-provenance-admission-external=true—Required. Task provenance must be protected by a controller-served or separately deployed webhook.
--workspace-class-use-admission-enabled=true—Required. The controller refuses to start without it.
--acp-workspace-dispatch-enabled—Lets agent Tasks actually request a workspace.
--agent-sandbox-enabled or --substrate-enabled—Picks the backend. Without one, workspace Tasks fail closed.
--substrate-direct-egress-enabledORKA_SUBSTRATE_DIRECT_EGRESS_ENABLEDRequired for native Substrate ACP admission. Acknowledges ateapi's --egress-gateway-address= configuration so worker NetworkPolicies see actual destinations. Defaults to false; suspension and cleanup still work.
--enable-fake-workspace-providerORKA_ENABLE_FAKE_WORKSPACE_PROVIDERDevelopment only — see below.

The source Helm chart enables both admission gates for harness-v2. It does not expose values for the provider API, workspace dispatch, or backend gates.

Task provenance protection is required because Orka stores workspace settlement state under reserved acp.workspace.orka.ai/ Task metadata. The provenance webhook prevents clients from forging that metadata. Use --task-provenance-admission-enabled for the controller-served webhook. With the dedicated admission runtime, install and verify the fail-closed webhook configuration before enabling --task-provenance-admission-external.

Upgrades need the CRDs applied by hand

Helm installs a chart's crds/ on first install and never updates them. Before enabling this on an existing cluster, apply the target chart's CRDs so the workspace.orka.ai schemas match the controller:

helm show crds '<chart>' | kubectl apply --server-side -f -

See Upgrading.

The fake adapter is for development only, and the release chart deliberately leaves its two CRDs out. Install them from a matching source checkout:

bin/kustomize build --load-restrictor LoadRestrictionsNone \
config/development/fake-workspace-provider | kubectl apply -f -

Who owns what​

ObjectScopeOwned byHolds
ExecutionWorkspaceProviderclusteroperatorThe adapter identity. For the in-tree adapter this is exactly controllerName: acp.workspace.orka.ai/runtime-pool.
RuntimeProviderConfigclusteroperatorWhich backend — agent-sandbox or substrate.
RuntimeWorkspaceProfilenamespacedoperatorBackend inputs: a Substrate profile names the infrastructure ActorTemplate and may set substrate.suspend; an agent-sandbox profile is empty unless the class allows suspension.
ExecutionWorkspaceClassnamespacedoperatorWhat users pick by name.
Task.spec.execution.workspace.classRef—userThe choice. Nothing else.

That split is the point: users name a class, and provider identity, backend parameters, pool implementation, and provider versions all stay with the operator. The older direct agent-sandbox and Substrate settings below still work during migration.

kubectl get executionworkspaceclass shows each class's lifecycle rules, so a person choosing a class can see whether their workspace is kept asleep (Suspend) or deleted (Delete) when the agent stops, how long it may sit idle, and how long it may exist:

$ kubectl get executionworkspaceclass
NAME MODE PROVIDER ON DETACH IDLE TIMEOUT MAX LIFETIME READY AGE
sandbox-session Interactive agent-sandbox Suspend 30m 24h True 2d
scratch Interactive agent-sandbox Delete 2h True 2d

-o wide adds the detach timeout. Helm does not update CRDs during an upgrade, so the columns appear once the CRDs from the new chart are applied (see Upgrading).

Who is allowed to use a class​

Selecting a class is an authorization decision, checked with a live Kubernetes SubjectAccessReview for verb use on that ExecutionWorkspaceClass. A denied or unavailable SAR denies the request.

Shipped ValidatingAdmissionPolicy resources enforce this at all times, even with the workspace gates off. With the provider API enabled, a TLS-backed webhook is required as well (--workspace-class-use-admission-enabled). How that webhook gets installed depends on how you installed Orka:

InstallWebhooksWhat you do
Helm, harness-v2task-workspace-class.harness-v2.orka.ai, tool-workspace-class.harness-v2.orka.aiNothing — the chart renders them against the release's own webhook Service and sets the flag. Requires webhooks.tls.existingSecret plus either webhooks.caBundle or webhooks.caInjectionAnnotations.
Kustomizetaskworkspaceclassuse.core.orka.ai, toolworkspaceclassuse.core.orka.aiApply config/orka-admission first, then config/orka-admission-webhooks once its README's readiness and TLS prerequisites are met.
Do not apply the Kustomize admission packages to a Helm release

You get two independent sets of validating webhooks for the same resources.

Binding a Task to a class​

spec:
execution:
workspace:
classRef:
name: my-workspace-class
reusePolicy: session # optional
workspaceSlot: default # optional
onDetach: Delete # optional

On admission the controller freezes the class identity, a hash of the profile, the lifecycle settings, and the effective detach action into the Task's immutable execution snapshot. Later edits to the class do not reach in-flight Tasks.

workspaceSlot composes with reusePolicy: none. Session reuse supports only the default slot and fails closed for any other value, until RuntimeSession controls become slot-scoped.

What happens when the agent detaches​

onDetachagent-sandboxSubstrate
DeleteWorksWorks
SuspendWorks, when the profile sets agentSandbox.suspendWorks with a DataOnly Substrate profile and the native provider pin

Delete is always executable.

Agent-sandbox suspension is cold, never a memory snapshot. The profile's agentSandbox.suspend freezes a durable workspace PVC shape (capacity, optional storageClassName and accessModes); the pool's SandboxClaim requests that PVC, which forces a cold start instead of adopting a warm sandbox. Suspending patches that exact Sandbox to operatingMode: Suspended so its Pod terminates while the PVC survives. Resume rotates the bootstrap material, refreshes the Sandbox blueprint, and returns the Sandbox to Running against the preserved volume.

Substrate DataOnly suspension uses verified native Tags, exact worker Pod termination, and fresh Actors for cold continuation. Its native lifecycle calls have no UID/version preconditions; Orka journals its own operation intents and requires fresh authenticated admission. Full-memory restore remains disabled. See Substrate workspaces for the trusted provider boundary, checkpoint export, restore authorization, and explicit recovery.

Expiry​

A suspend-capable class must set idleTimeout or maxLifetime. Without one, class readiness and Task binding fail closed.

SettingWhereEffect
idleTimeoutclass lifecycleIdle suspended workspaces expire; idle Ready workspaces take the class default action.
maxLifetimeclass lifecycleHard cleanup bound, always.
retention.maxSuspendedWorkspacesRuntimeWorkspaceProfileCaps concurrently suspended workspaces per class and namespace. Rejected at admission, retried at settlement with the frozen Suspend action preserved. This is a cap, not an expiry.

A queued continuation may take a still-Ready workspace directly. Deletion policies that retain data past workspace deletion are rejected.

ADRs 0026–0030 carry the full contract.

Agent Sandbox controller settings​

Workspace-provider-backed ACP RuntimeSession dispatch requires --acp-workspace-dispatch-enabled plus the matching provider flag (--agent-sandbox-enabled or --substrate-enabled); with either unset, Task.spec.execution.workspace agent Tasks fail closed. The Substrate backend also uses --substrate-api-*, --substrate-router-url, and --substrate-actor-dns-suffix. Native ACP additionally requires --substrate-direct-egress-enabled, available through Helm as controller.substrate.directEgressEnabled. See Substrate setup for the provider configuration this acknowledges. The agent-sandbox router, template, timeout, and cleanup settings below belong to the earlier worker-path prototype and are not used by the ACP RuntimePool backend, which renders its own sandbox templates:

FlagEnvironment variableHelm valueDefault
--agent-sandbox-enabledORKA_AGENT_SANDBOX_ENABLEDcontroller.agentSandbox.enabledfalse
--agent-sandbox-router-urlORKA_AGENT_SANDBOX_ROUTER_URLcontroller.agentSandbox.routerUrlempty
--agent-sandbox-default-templateORKA_AGENT_SANDBOX_DEFAULT_TEMPLATEcontroller.agentSandbox.defaultTemplateempty
--agent-sandbox-warm-pool-policyORKA_AGENT_SANDBOX_WARM_POOL_POLICYcontroller.agentSandbox.warmPoolPolicydisabled
--agent-sandbox-namespace-strategyORKA_AGENT_SANDBOX_NAMESPACE_STRATEGYcontroller.agentSandbox.namespaceStrategytask
--agent-sandbox-claim-timeoutORKA_AGENT_SANDBOX_CLAIM_TIMEOUTcontroller.agentSandbox.claimTimeout2m
--agent-sandbox-command-timeoutORKA_AGENT_SANDBOX_COMMAND_TIMEOUTcontroller.agentSandbox.commandTimeout30m
--agent-sandbox-cleanup-policyORKA_AGENT_SANDBOX_CLEANUP_POLICYcontroller.agentSandbox.cleanupPolicydelete

Supported values are disabled or template for the legacy warm pool policy setting, task or controller for namespace strategy, and delete or retain for the legacy cleanup policy. ACP ignores this cleanup default and rejects Task cleanupPolicy: retain. task defaults sandbox claims to the Task namespace; controller defaults them to the controller namespace when discoverable, and explicit templateRef.namespace values are honored as the claim/warm-pool namespace. See Agent Sandbox Workspaces for what the ACP-backed provider does today and the invariants it holds.

Any future ACP-backed integration will need a separately reviewed identity and RBAC design. Do not grant these permissions to managed ACP RuntimePods; they intentionally run without Kubernetes service-account tokens or Kubernetes RBAC.

External API OIDC authentication​

ServiceAccount bearer token authentication is always available. To allow external callers such as GitHub Actions to authenticate directly with OIDC JWTs, configure issuer, audience, an explicit subject allowlist, and the namespace assigned to OIDC callers:

--oidc-issuer=https://token.actions.githubusercontent.com
--oidc-audience=orka-ci
--oidc-allowed-subjects=repo:my-org/my-repo:ref:refs/heads/main
--oidc-namespace=ci

The same settings can be supplied with environment variables:

ORKA_OIDC_ISSUER=https://token.actions.githubusercontent.com
ORKA_OIDC_AUDIENCE=orka-ci
ORKA_OIDC_ALLOWED_SUBJECTS=repo:my-org/my-repo:ref:refs/heads/main
ORKA_OIDC_NAMESPACE=ci
# Optional; when omitted, Orka discovers the JWKS URL from the issuer metadata.
ORKA_OIDC_JWKS_URL=https://token.actions.githubusercontent.com/.well-known/jwks

OIDC validation requires RS256-signed JWTs with matching iss and aud, valid time claims, a non-empty sub, and a sub value that matches --oidc-allowed-subjects. Wildcards * and ? are supported in allowlist patterns; use the narrowest GitHub Actions subject for the trusted repository, branch, environment, or workflow. Authorized OIDC callers are bound to --oidc-namespace (or default when omitted) so namespace isolation can reject requests for other namespaces. When an OIDC-authenticated caller creates a Task, Orka records the verified identity in spec.requestedBy. Clients cannot set requestedBy themselves.

External API context-token authentication​

Orka can also authenticate external API requests with generic transaction/context tokens. The built-in transaction-token profile validates RS256-signed JWTs with JOSE header typ: txntoken+jwt, matching iss and aud, valid time claims, a non-empty sub, and the required transaction-token claims iat, txn, scope, and req_wl.

For the strict profile contract, see Transaction Tokens. For the breaking configuration change, see the migration guide.

Enable the profile by configuring the profile, issuer, and audience:

--context-token-profile=transaction-token
--context-token-issuer=https://issuer.example.com
--context-token-audience=orka-api

The same settings can be supplied with environment variables:

ORKA_CONTEXT_TOKEN_PROFILE=transaction-token
ORKA_CONTEXT_TOKEN_ISSUER=https://issuer.example.com
ORKA_CONTEXT_TOKEN_AUDIENCE=orka-api
# Optional for transaction-token; when omitted, Orka uses <issuer>/.well-known/jwks.json.
ORKA_CONTEXT_TOKEN_JWKS_URL=https://issuer.example.com/.well-known/jwks.json

By default, the transaction-token profile reads raw transaction tokens from the Txn-Token header:

curl -H "Txn-Token: $TXN_TOKEN" https://orka.example.com/api/v1/tasks

To customize token locations, set --context-token-headers or ORKA_CONTEXT_TOKEN_HEADERS to a comma-separated list. Use Header for raw token headers and Header:Scheme for scheme-prefixed headers. For example, keep the default Txn-Token header and explicitly opt in to Authorization: Bearer context-token support:

--context-token-headers=Txn-Token,Authorization:Bearer

Authorization: Bearer remains the default location for Kubernetes ServiceAccount and OIDC JWT authentication. Context-token bearer authentication is only attempted when Authorization:Bearer is explicitly configured and the bearer JWT has typ: txntoken+jwt; other bearer tokens continue through the standard OIDC or Kubernetes TokenReview flow. When an external context-token caller creates a Task, Orka records the verified subject and issuer in immutable spec.requestedBy and records safe transaction metadata in immutable spec.transaction, transaction labels, and transaction annotations. Clients cannot set requestedBy or transaction themselves.

Optional authorization is controlled by --context-token-authz-mode / ORKA_CONTEXT_TOKEN_AUTHZ_MODE. In audit mode, Orka logs safe authorization failures and allows the request. In enforce mode, Orka rejects context-token callers that lack the configured operation scope or violate signed tctx constraints. Task creation can be constrained by tctx.namespace, tctx.taskType, tctx.agent, tctx.allowedAgents, workspace tctx.repo/tctx.branch/tctx.ref, and tctx.allowedTools. Chat, OpenAI-compatible, and Anthropic-compatible model calls require the provider-use scope (default orka:providers:use) and honor tctx.namespace, tctx.provider, tctx.allowedProviders, tctx.model, and tctx.allowedModels. When Orka-managed server-side tools are exposed to those endpoints, they also require the tool-use scope (default orka:tools:use) and honor tctx.allowedTools. Security scan read/list/get endpoints require the security-read scope (default orka:security:read), and security scan create/update/delete and mutation endpoints require the security-write scope (default orka:security:write). Repository monitor read endpoints require the monitor-read scope (default orka:monitors:read), monitor create/update/delete endpoints require the monitor-write scope (default orka:monitors:write), and manual monitor runs require the monitor-operate scope (default orka:monitors:operate). Repository monitor access can also be constrained by tctx.namespace, tctx.repo, tctx.branch, tctx.agent, and tctx.allowedAgents. The raw TxToken is never logged or persisted in Task specs/status.

Transaction-token TTS exchange and propagation​

Configure --context-token-tts-endpoint / ORKA_CONTEXT_TOKEN_TTS_ENDPOINT when workers should exchange a mounted subject token for child or outbound replacement TxTokens. Delegation tools require ORKA_CONTEXT_TOKEN_SUBJECT_TOKEN_FILE and ORKA_CONTEXT_TOKEN_CHILD_SCOPE; HTTP Tool calls can use ORKA_CONTEXT_TOKEN_OUTBOUND_SCOPE or fall back to the current transaction scope. Child scopes are fail-closed: Orka rejects a requested child scope that is not already present in the parent transaction scopes before it creates the child Task.

Successful delegation exchanges store the raw child TxToken only in an owner-referenced Kubernetes Secret and annotate the child Task with the Secret name. The controller mounts that Secret into the child worker and sets ORKA_TRANSACTION_TOKEN_FILE / ORKA_CONTEXT_TOKEN_SUBJECT_TOKEN_FILE so deeper delegation and downstream Tool calls can continue the same transaction with configured child/outbound scopes.

Task provenance admission hardening​

The REST API rejects client-supplied requestedBy and transaction fields and stamps verified provenance itself. To also protect direct Kubernetes Task CRD writes, enable the optional validating admission webhook:

--task-provenance-admission-enabled=true

The webhook denies untrusted CREATE or UPDATE requests that set or modify Orka-managed provenance fields: spec.requestedBy, spec.transaction, orka.ai/transaction-* labels/annotations, orka.ai/context-token-profile, and the child token Secret annotation. By default, trusted writers are the Orka controller ServiceAccount usernames in the controller namespace and the configured AI and vendor worker ServiceAccount names in the target Task namespace; override them with --task-provenance-admission-trusted-users and --task-provenance-admission-trusted-service-accounts.

Worker-created coordination children must name the caller's active Task as parent, verified through the request's Pod identity and the live Pod, Job, and Task ownership chain. A worker-created Task may omit sessionRef or inherit its parent's reference unchanged, including any history cutoff. Workers cannot introduce arbitrary sessions through a new child or an ownerless Task. Trusted controller writers can establish separately authorized session references.

The webhook also makes coordination parent names and the parent Task controller-owner identity immutable after creation, including for trusted workers and controllers. Kubernetes cleanup controllers may remove ownership when orphaning a dependent; that Task then loses access through the removed parent. Internal cross-task transcript searches and coordination messages require protection through manager-hosted admission (--task-provenance-admission-enabled=true) or a separately deployed webhook (--task-provenance-admission-external=true). With both flags disabled, internal transcript search is limited to the caller's own session and coordination message requests are denied.

Enabling either flag asserts that all retained Task ancestry was created under these checks. Admission does not validate old edges retroactively, and an unchanged update does not attest them. Before first enabling ancestry trust, or upgrading from a validator that did not authenticate coordination parents, follow the Task cleanup requirement below.

Internal transcript search applies sessionRef.maxMessages and sessionRef.throughMessageId before matching messages or limiting results. When multiple Tasks reference the same session, search uses the intersection of their permitted history windows. An unbounded reference cannot widen a bounded one, and a missing cutoff message yields no matches for that session.

How admission is deployed depends on the installation method. Helm releases install and enable Task-provenance admission automatically: the chart renders task-provenance.<mode>.orka.ai with failurePolicy: Fail against the release-local controller webhook Service and runs the controller with --task-provenance-admission-enabled=true, trusting the release controller identity. For Kustomize installations, admission deployment is opt-in and served by the dedicated admission runtime, not the controller manager: install config/orka-admission (Deployment, Service, NetworkPolicy, and RBAC for the admission runtime), then apply config/orka-admission-webhooks — which includes taskprovenance.core.orka.ai with failurePolicy: Fail — only after the readiness, TLS Secret, and CA-injection prerequisites in config/orka-admission-webhooks/README.md are met and the trusted identities embedded in validating_webhook.yaml match the admission-runtime arguments.

The first config/acp-production workload wave leaves ancestry trust disabled. Enable --task-provenance-admission-external=true in the controller's post-admission configuration only after the separate webhook is installed and rejects unauthorized Task provenance changes. The flag enables ancestry trust for both internal worker API and brokered MCP coordination policy. Keep it disabled if the webhook wave is omitted, and disable it and complete the controller rollout before removing the webhook.

Upgrading internal worker authorization​

If existing Tasks were created without authenticated ancestry admission, prevent new Task creation in each affected watched namespace throughout the cutover. Drain all Tasks under the old controller, including ACP and remote-runtime agent Tasks. Preserve required outputs outside Task cleanup, then delete every pre-existing Task in those namespaces and wait for its finalizers to finish. Retained terminal Tasks must also be removed because their session references participate in coordination searches.

Install the updated fail-closed admission policy before enabling either provenance flag and allowing Task creation again. An unchanged update or retained labels and owner references do not attest legacy ancestry. Helm enables the flag automatically, so complete this cleanup before upgrading that release. Routine restarts with continuously enforced ancestry admission do not require this cleanup.

Before upgrading from a controller that does not record Task.status.jobUID, pause Task producers and drain all Job-backed Tasks while the old controller is still running. Wait until their attempts are terminal and their results and artifacts are stored. Apply the CRDs from the exact target chart, upgrade the controller and admission components, then resume Task producers. A rolling upgrade with active legacy Job-backed Tasks is not supported.

The new authorization checks require the controller-recorded Job UID. They reject workers whose Tasks have only status.jobName, including otherwise valid Pods from the previous controller. Orka does not backfill UIDs by looking up Job names: namespace Job creators can replace those Jobs. Do not patch missing UIDs from a name lookup; finish the attempt before upgrading or explicitly submit a new Task after the upgrade.

Built-in harness v1 artifact uploads also require the wrapper's Kubernetes workload identity. Configure the bound wrapper endpoint as a Kubernetes Service in the auth Secret's namespace. The uploading Pod must belong to a live ReplicaSet and Deployment selected by that Service. A worker Pod with a copy of the wrapper bearer does not receive artifact access.

Prometheus metrics​

Orka registers the following Prometheus metrics on the controller-runtime registry. The metrics endpoint is disabled by default (--metrics-bind-address=0); enable it by setting an explicit bind address, for example:

--metrics-bind-address=:8443 # HTTPS (default when --metrics-secure=true)
--metrics-bind-address=:8080 # HTTP, with --metrics-secure=false

Scrape configuration (for example a Prometheus Operator ServiceMonitor) is not shipped with the chart; point your monitoring stack at the metrics port directly.

MetricTypeLabelsDescription
orka_api_requests_totalCounterendpoint, method, statusTotal API requests (status bucketed as 2xx/4xx/5xx)
orka_api_request_duration_secondsHistogramendpoint, methodAPI request latency in seconds
orka_skills_loaded_totalCounterskill, namespaceSkills loaded by namespace and name
orka_store_db_size_bytesGauge—Size of the SQLite database file in bytes
orka_context_token_auth_totalCounterprofile, resultContext-token authentication attempts
orka_context_token_authorization_totalCounteraction, result, reasonContext-token authorization decisions (allow/deny/audit)
orka_context_token_tts_exchange_totalCounterresult, reasontransaction-token TTS token-exchange attempts
orka_context_token_tts_exchange_duration_secondsHistogramresult, reasontransaction-token TTS token-exchange latency in seconds
orka_token_exchange_totalCounteradapter, grant_class, result, reasonDirect and transaction-token OAuth exchange attempts
orka_token_exchange_duration_secondsHistogramadapter, grant_class, result, reasonOAuth exchange latency in seconds
orka_acp_runtime_pool_desired_replicasGaugenamespace, runtime_poolDesired RuntimePool replicas
orka_acp_runtime_pool_ready_replicasGaugenamespace, runtime_poolAuthoritatively selected Ready RuntimePool replicas
orka_acp_runtime_pool_sessions_activeGaugenamespace, runtime_poolAuthenticated resident RuntimeSession count
orka_acp_runtime_pool_prompts_in_flightGaugenamespace, runtime_poolAuthenticated active prompt count
orka_acp_runtime_pool_queued_tasksGaugenamespace, runtime_poolDurable unsatisfied Task demand assigned to the pool
orka_acp_runtime_pool_admission_stateGaugenamespace, runtime_pool, stateOne-hot authoritative admission state (unknown, closed, accepting, draining, or ambiguous)
orka_acp_runtime_pool_scale_to_zero_totalCounternamespace, runtime_poolCompleted RuntimePool scale-to-zero transitions

Context-token metrics are described in more detail in Transaction Tokens. All context-token labels use low-cardinality values only.

OpenTelemetry telemetry​

Orka supports opt-in OpenTelemetry traces and GenAI metrics for debugging, performance analysis, and backend cost/latency dashboards. Telemetry is disabled by default and uses OpenTelemetry no-op providers until enabled.

Enabling telemetry​

Add the --enable-telemetry flag to the controller and configure an OTLP collector endpoint. The legacy --enable-tracing alias enables the same traces and metrics:

args:
- --enable-telemetry
env:
- name: OTEL_EXPORTER_OTLP_ENDPOINT
value: "http://jaeger-collector.observability.svc:4317"
- name: OTEL_EXPORTER_OTLP_INSECURE
value: "true"
Flag / Environment VariableDefaultDescription
--enable-telemetry / --enable-tracingfalseEnable OpenTelemetry traces and metrics
OTEL_EXPORTER_OTLP_ENDPOINTSDK default localhost:4317OTLP collector endpoint for traces and metrics
OTEL_EXPORTER_OTLP_TRACES_ENDPOINTunsetTrace-specific OTLP endpoint
OTEL_EXPORTER_OTLP_METRICS_ENDPOINTunsetMetrics-specific OTLP endpoint
OTEL_EXPORTER_OTLP_PROTOCOLSDK default gRPCSet to http/protobuf for OTLP/HTTP collectors
OTEL_EXPORTER_OTLP_TRACES_PROTOCOL / OTEL_EXPORTER_OTLP_METRICS_PROTOCOLunsetSignal-specific exporter protocol overrides
OTEL_EXPORTER_OTLP_INSECURE and signal-specific insecure varsSDK defaultDisable TLS for in-cluster/dev collectors that require it
OTEL_TRACES_SAMPLER / OTEL_TRACES_SAMPLER_ARGSDK defaultStandard OpenTelemetry sampler configuration

Controller-local defaults such as localhost:4317 are valid only for the controller process. AI worker Jobs receive telemetry enablement only when the controller has a non-loopback, worker-reachable OTLP endpoint. The controller copies non-secret OTLP endpoint/protocol/insecure/timeout/compression settings to AI worker Pods and intentionally does not copy OTLP headers, certificate or client-key env vars, OTEL_RESOURCE_ATTRIBUTES, or baggage.

ACP attempt, RuntimeSession, and publication spans use the controller exporter. With controller telemetry enabled and a reachable trace endpoint, managed RuntimePool supervisors receive non-secret endpoint, protocol, insecure and compression settings. Supervisor endpoints require http:// or https:// URLs; unsupported SDK settings disable export. Authenticated v2 operations continue the current request's W3C trace context. Provider children receive no tracing settings. Collector routing is operator-owned; enabling telemetry does not change runtime network policies. See supervisor configuration and routing.

Instrumented components​

TracerSpanAttributes
orka.apiHTTP/API middleware spansHTTP request/route/status metadata
orka.chatchat.request, chat.tool_loop.iterationsession metadata; chat.iteration, orka.tenant, requested model, tool-call count
orka.workertask.runorka.task.id, namespace, and agent name when known
orka.acpacp.prompt, acp.session.create, acp.session.continue, acp.publication.reconcileorka.task.id, namespace, attempt/prompt identity, RuntimePool/RuntimeSession identity, prompt/session outcome, publication identity, and agent name when known
orka.acp.supervisoracp.supervisor.* authenticated runtime operationsTask UID/attempt, prompt/operation IDs and runtime pool/session identity; no request content
orka.agentagent.stepiteration, requested model/provider, tool-call count, Orka task metadata
orka.gen_aichat {model}gen_ai.* provider/model/token metadata and error.type
orka.gen_aiexecute_tool {tool.name}gen_ai.tool.*, orka.tool.name, orka.tool.kind, orka.tool.result.size_bytes, parent/child task fields for delegation
orka.controllertask.reconciletask name, namespace, type, and propagated trace context

Use orka.task.id to find a Task trace, orka.tool.name to find specific tool executions, and orka.parent_task.id / orka.child_task.id to follow delegated children. Tool spans do not include raw arguments or result bodies.

Example: Jaeger setup​

These commands assume a Helm install named orka

That makes the Deployment orka-controller. If you installed with kubectl apply -f .../deploy/orka.yaml it is orka-controller-manager instead. See Troubleshooting.

# Deploy Jaeger all-in-one (development only)
kubectl create namespace observability
kubectl apply -n observability -f https://raw.githubusercontent.com/jaegertracing/jaeger-operator/main/examples/simplest.yaml

# Configure the controller
kubectl -n orka-system set env deployment/orka-controller \
OTEL_EXPORTER_OTLP_ENDPOINT=http://jaeger-collector.observability.svc:4317

Context engineering best practices​

Research on LLM agent context files (arxiv 2602.11988) shows that verbose context hurts more than it helps: LLM-generated context files reduce task success rates by 0.5–2% while increasing inference costs by 20–23%. Even developer-written files yield only marginal improvements (~4%) with similar cost increases. The guidelines below translate these findings into practical advice for Orka's systemPrompt, Skill, and Agent configuration.

Writing effective system prompts​

Keep Agent systemPrompt content minimal and requirement-focused:

  • Include only: tooling commands (build/test/lint invocations), non-discoverable gotchas (e.g., "provider secret key defaults to api-key"), and hard constraints the agent cannot infer from source code.
  • Avoid codebase overviews — agents discover project structure efficiently on their own through file listing and search tools. Overviews add tokens without improving navigation speed.
  • Don't duplicate information already present in website/docs/, README, or inline code comments. Redundant instructions increase reasoning token usage (14–22% more) without improving outcomes.

Writing effective Skills​

Skills are prepended to the system prompt on every LLM call for every task that uses the parent Agent. Each Skill directly increases per-request token cost.

  • Keep Skill content concise and action-oriented — write instructions ("run make lint-fix after changes"), not descriptions ("this project uses a Makefile-based build system").
  • Split large Skills so Agents only reference the ones they need. A research Agent doesn't need a coding-standards Skill.
  • Regularly audit Skill content and remove instructions the agent follows by default.

Monitoring recommendations​

  • With GenAI telemetry enabled, track gen_ai.client.operation.duration and gen_ai.client.token.usage to compare model-call latency and token use.
  • A/B test agent performance with and without specific systemPrompt or Skill content to validate that each addition provides measurable benefit.
  • Well-documented repositories benefit least from additional context; focus context engineering effort on repos with limited existing documentation.