Skip to main content

Operating generic gateways

Generic gateways are enabled by default. They use the controller SQLite database for normalized ingress events, Session transcript provenance, and outbound delivery state.

External adapters​

The experimental A2A adapter lets A2A clients call Orka Agents through the gateway. It runs separately from Orka and supports a text-only A2A subset.

See its setup guide and compatibility notes for supported Orka versions, native-Agent prerequisites, and conformance limits.

Default service levels and bounds​

The documented local reference target is p95 durable ingress acknowledgement below 500 ms for a 1,000-event burst with 100 active Sessions. This is an admission SLO, not an end-to-end model response SLO.

Defaults:

SettingDefault
Pending events per Session100
Retained operational event records per Gateway1,000
Retained rejected audit records per Gateway250
Event/delivery expiry24h
Delivery timeout15s
Delivery attempts10
Distinct accepted interim messages per Task10 (--gateway-interim-messages-per-task)
Interim message text16 KiB UTF-8
Ingress and terminal text64 KiB
Terminal retention720h (30 days)
Claim lease1m
Default persistent volume request1Gi

Controller flags use the --gateway-* prefix. Helm values are under controller.gateway. Gateway dispatch, delivery, and retention maintenance run only on the elected controller leader; SQLite claims preserve crash recovery while leader election makes namespace task-capacity checks replica-safe. Gateway ingress is enabled by default, so the Helm chart also enables the SQLite PVC by default. Helm rejects controller.gateway.enabled=true with store.persistence.enabled=false unless controller.gateway.allowEphemeralStore=true is explicitly set for disposable development; acknowledged events on an ephemeral store are not durable across Pod replacement. For offline GitOps rendering, apply the bundled Task/Gateway CRDs separately and set controller.gateway.crdsReadyOverride: true; live Helm installs normally discover the CRDs with lookup.

Readiness​

A Gateway is Ready only when:

  • its GatewayClass is accepted;
  • exactly one safe HTTPS endpoint or TLS-authenticated selector-backed same-namespace Service resolves;
  • separate inbound/outbound Secrets exist, opt in, and bind to the Gateway;
  • the outbound Secret binds to the exact resolved HTTPS endpoint;
  • authenticated health and capability probes succeed;
  • observed capabilities satisfy the GatewayClass requirements.

Agent runtime warmth does not affect Gateway readiness. Accepted events remain durable while downstream execution is temporarily unavailable.

serviceRef TLS and private CAs​

A serviceRef resolves to https://<service>.<namespace>.svc:<port>. The adapter certificate must include <service>.<namespace>.svc as a DNS subject alternative name; an IP address or a different external hostname does not satisfy controller hostname verification.

For a private CA, mount the CA certificate into the controller and either install it in the container's system trust store or set SSL_CERT_DIR to the mounted CA directory. Do not disable certificate verification. Roll the controller after changing its trust configuration, and wait for the rollout before expecting the Gateway to become Ready. For a Helm release named orka in orka-system (a kubectl apply install names the Deployment orka-controller-manager instead — see Troubleshooting):

kubectl -n orka-system rollout restart deployment/orka-controller
kubectl -n orka-system rollout status deployment/orka-controller

Agent execution​

GatewayBinding.spec.agentRef supports both native AI Agents and runtime-backed Agents. At Task creation, an Agent without spec.runtime produces a type: ai Task using the Agent's model/provider configuration and the AI worker. An Agent with spec.runtime continues to produce a type: agent Task. taskDefaults.agentRuntimeMaxTurns is runtime-only and makes a binding to a native AI Agent not ready; dispatch also rejects it if an Agent edit races readiness reconciliation.

Admission pins the Agent UID, not its generation or execution kind. Agent edits before Task creation can therefore change the selected execution kind. Once the Task exists, crash recovery adopts that Task with its immutable kind instead of selecting again from the edited Agent. Frozen external-runtime tool policy remains authoritative and cannot be discarded by switching to native AI.

Both paths consume the canonical Session transcript through the admitted user message, without copying external text into the Task CR or adding the current prompt twice. Native AI startup fails closed when a required transcript has no non-empty final user turn. Gateway terminal projection—not generic Task finalization—appends the canonical assistant message, creates the delivery, and releases the Session lock.

Interim messages​

An adapter may advertise optional interimDelivery: true under orka.gateway.v1. Only then can a Running gateway Task enqueue a bounded kind: message delivery. Absent/false capability means final/error-only operation; no message is enqueued or sent. This adds a kind, not delivery statuses: existing final/error meaning and delivered/retryableError/nonRetryableError receipts are unchanged.

Interim text is kept in delivery records, not in the Task CR, final result, or canonical Session transcript. It does not complete the event, annotate the Task as delivered, or unlock the Session. Use the gateway delivery view for progress; normal terminal projection still owns completion and history.

The controller limit is --gateway-interim-messages-per-task=10. Supply a positive value to change it; the setting is not a per-worker routing option. Every distinct accepted message counts for that Task's lifetime, including failed/expired messages. An exact request replay returns the retained receipt without another enqueue or quota charge. Text is limited to 16 KiB before sanitization, without silent truncation; terminal text remains bounded at 64 KiB.

Replying while the agent continues​

The production native worker and ACP broker provide gateway-only reply_in_conversation:

{"content":"A bounded intermediate update"}

content is the only model argument. Its schema requires a string with minLength: 1, maxLength: 16384, and no additional properties. Execute also rejects invalid Unicode, unknown/duplicate fields, empty or sanitized-empty text, and content over 16384 UTF-8 bytes. There is no destination, request ID, count override, or approval argument. Success returns only deliveryID, current durable status, and created inside the tool success envelope. A receipt means durable acceptance, not completed delivery, and the agent continues working normally.

The controller enables the tool only after proving exact admitted event/Task UID origin. Names, environment flags, prompt text, and provenance metadata cannot grant it. Ordinary AI/ACP Tasks, delegated children, container Tasks, and compatibility-proxy callers do not gain the tool. Explicit tool denials, closed allowlists, and transaction scopes remain authoritative. Built-in ACP providers preserve their native default tools. An external runtime must explicitly opt in through its registered, conformed profile with exact policy agreement; existing profiles without reply remain final-only. Already-frozen sessions are not upgraded.

For native AI, omitted Agent.spec.tools and Task.spec.ai.tools leave automatic reply eligibility unchanged; an explicit tools: [] in either field denies reply even with a proven gateway origin. Nonempty lists must permit reply, and an Agent reply_in_conversation entry with enabled: false remains a denial. Typed Kubernetes JSON writes preserve explicit empty lists. Older writes may already have collapsed an empty list into omission: persisted omission cannot reveal that original intent and is not reinterpreted during upgrade. Where denial was intended, reapply the explicit policy using an updated writer before enabling gateway replies; older writers can still erase empty lists.

The host—not the model—supplies a stable operation identity. The tool checks a controller-backed, replay-aware budget before enqueue; final atomic admission still enforces the configured lifetime cap against concurrent calls. Retained receipts can replay at exhaustion without another charge. Distinct logical calls with identical text are distinct messages. Retries must preserve operation identity; regeneration after a restart does not guarantee exactly-once delivery. Known unsupported/capacity/conflict rejections are safe model-visible errors. An uncertain send must not be treated as definite rejection or retried by inventing a new call; ACP retains its consequential OutcomeUnknown fence as well as gateway receipts.

Native and ACP authorization​

Native workers authenticate durable origin through GET /internal/v1/tasks/{namespace}/{taskName}/gateway-messages/origin, which returns only the exact Task UID. They read GET .../gateway-messages/budget?requestID=... for accepted, limit, and requestExists, then use POST .../gateway-messages with content and a host-derived request ID. All routes retain authentic current Pod/Job/Task UID authorization and revocation checks. The production client uses the projected token file, not an environment token override. New admission returns HTTP 202; replay returns HTTP 200. These are execution-host APIs, not operator/admin send endpoints.

Origin is separate from adapter availability: a proven gateway execution retains the tool through readiness/capability withdrawal, but new budget/admission requests still fail the current gate. Missing interim capability returns interim_delivery_unsupported (HTTP 409) without enqueueing. An unproven origin never enables the tool. The client retries origin 503 responses within a bounded startup window; a persistent transient origin-service failure omits this optional tool while normal work continues. Invalid identity is still denied. There is no late tool upgrade in that execution.

ACP does not use those native HTTP routes or create a Job. Its sender uses the signed broker call's active prompt/session guard and the authorized task-data transaction, independently for budget and enqueue. Frozen policy, exact Task identity, prompt authority, and durable gateway origin all remain required. Live preparation occurs outside the SQLite writer, with durable admission and mutation fences inside it. No approval or terminal-lifecycle behavior is added.

Task and Session access​

Gateway-created Tasks remain ordinary Orka Task objects for controller execution, but their CRs contain no external message text; the prompt is loaded from the bounded, task-owned Session transcript. Public Task list/get/log/result/event/trace/fork surfaces require both the ordinary Task permission and gateway-read authorization for the owning Gateway. Destructive Task actions and approval decisions additionally require gateway-operate authorization. The default Helm and Kustomize installs also create a fail-closed ValidatingAdmissionPolicy that permits direct Kubernetes create/update/delete of gateway-owned Tasks only from the owning Orka controller or trusted worker ServiceAccounts. Namespace-isolated Helm releases scope the policy by the immutable Gateway namespace encoded in requestedBy.issuer, so multiple releases do not deny one another and coordinated workers can create inherited child Tasks. Canonical gateway Sessions are hidden from generic Session, Session-event, transcript-search, and chat-loading surfaces; gateway event/delivery APIs are the supported operator view.

Secret rotation​

Create or update Secret data without changing the configured key. Kubernetes Secret watch events trigger reconciliation; status records only the observed resourceVersion.

Token values belong only in the Secret

Never put them in Gateway metadata, status, logs, Tasks, or support bundles. Status records the Secret's resourceVersion and nothing else, and support bundles are routinely shared.

When an endpoint changes, update the outbound Secret's gateway.orka.ai/adapter-endpoint annotation to the exact new resolved HTTPS endpoint. The Gateway remains not ready until the binding matches and the authenticated probe succeeds.

Dead letters and recovery​

Inspect event and delivery state in the dashboard or with:

orka gateway events list --state DeadLettered,Expired
orka gateway deliveries list --state DeadLettered,Failed,Expired
orka gateway deliveries retry '<delivery-id>'

Manual retry preserves the stable delivery/idempotency ID, resets the bounded attempt window, extends expiry by 24 hours, and increments manualRetryCount. Expired ingress events are not automatically replayed because their original sender/context authorization may no longer be valid.

Within an event, live interim predecessors drain in order before later messages and final/error. Permanent failure, exhausted attempts, or expiry abandons a message so terminal delivery can proceed. Capability withdrawal is rechecked before sending and prevents a queued interim post. An earlier message cannot be manually revived once any later delivery has started, including Sending, uncertain, or subsequently expired/abandoned attempts. Do not regenerate IDs to work around this fence. Terminal manual retry is unchanged. Replaying an old delivered message ID after terminal must return its original provider correlation without sending again.

Backup and restore​

Gateway state lives in two places, and a backup of one alone is not a backup.

WhereWhat is in it
Controller SQLite database (on the PVC)Gateway Sessions, normalized events, delivery rows, event-to-Task correlation
KubernetesGateway CRDs, referenced Secrets, Task CRs
Back up both, and keep credentials out of the SQLite archive

A disaster-recovery backup needs the SQLite volume and the corresponding cluster objects. Credentials belong in the cluster Secret backup only.

SQLite runs in WAL mode, which means a committed record may still be sitting in orka.db-wal rather than in orka.db.

Copying orka.db alone from a running controller silently loses committed work

There is no error. The copy simply comes back missing whatever was still in the WAL.

Use one of these consistency-safe approaches:

  • take an atomic CSI VolumeSnapshot of the whole persistent volume, including the database, WAL, and shared-memory files; quiescing writes first is still preferred;
  • use SQLite's online backup API against the live database; or
  • pause adapter ingress, stop every controller replica that can write the store, checkpoint/truncate the WAL from a maintenance process, and then copy the main database file.

For a release or disaster-recovery backup:

  1. Record the Orka image/chart version, Gateway CRD manifests, controller gateway flags, and adapter versions/capabilities.
  2. Pause external ingress and wait for in-flight API requests to finish. Record queued, sending, retry-scheduled, and dead-letter counts.
  3. Create a consistent SQLite/PVC snapshot and a cluster backup containing GatewayClasses, Gateways, GatewayBindings, Tasks, referenced Agents, and referenced Secrets.
  4. Verify the SQLite copy with PRAGMA integrity_check in an isolated location and keep the backup immutable.
  5. Resume ingress only after the snapshot and cluster-object backup both succeed.

To restore, stop gateway writers, restore the SQLite files and matching Kubernetes objects, then start one controller replica first. Confirm CRDs are Established, the store opens successfully, Gateways return Ready, terminal deliveries remain terminal, and queued events/due deliveries become claimable before restoring normal replica count and ingress. Claims that were active at backup time become eligible only after their recorded lease expires.

Deterministic Task and delivery IDs prevent restored work from receiving new identities. They cannot prove whether a provider accepted a request immediately before the snapshot, so a conforming adapter must deduplicate any replay by the original delivery/idempotency ID.

Never "repair" a restore by deleting or regenerating delivery IDs

Those IDs are the only thing stopping a replayed delivery from being processed twice. A new ID turns a safe duplicate into a second real side effect.

Upgrade compatibility and version skew​

The adapter wire contract is exact-versioned. The current controller accepts only orka.gateway.v1; adapterVersion is informational and capabilities are readiness inputs, not a version-negotiation mechanism. Unknown JSON fields are rejected. There is no implied N-1 or N+1 adapter compatibility: a controller and adapter may be rolled independently only while both continue to speak exactly orka.gateway.v1 and the adapter still advertises every capability required by its GatewayClass.

The interim-delivery extension requires controller-first rollout. New controllers accept old adapters without the capability. Older strict controllers reject the interimDelivery field: an unchanged orka.gateway.v1 discriminator does not imply reverse-skew compatibility. The reference adapter now advertises true by default; use --interim-delivery=false to omit the field for legacy controllers.

Before an older-controller rollback, pause ingress/new interim work, drain or abandon outstanding message deliveries with the new controller, and disable adapter advertisement. Verify observed capabilities reflect the opt-out before replacing the controller. Disabling the flag alone does not make retained pending message rows safe for an old dispatcher. No schema change accompanies this feature, but normal behavioral, CRD, and database rollback checks still apply.

Use this rollout order:

  1. Take the consistency backup above and export the currently installed Gateway CRDs.
  2. Apply the target release's CRDs first and wait for each CRD to become Established. Do not remove the currently served/storage version during a rolling upgrade.
  3. Roll the controller and verify store startup, API health, Gateway readiness, queue depth, and dead-letter rate.
  4. Run the gateway conformance CLI against each target adapter build, then roll adapters one at a time.
  5. Confirm observed adapter name/version/capabilities and perform one idempotent test delivery before completing the rollout.

Run conformance from a network location that can reach the target adapter. Set SSL_CERT_FILE when the adapter uses a private CA that is not already trusted by the host:

SSL_CERT_FILE=/path/to/ca.crt \
ORKA_GATEWAY_BEARER_TOKEN='<outbound bearer token>' \
go run ./cmd/orka-gateway-conformance \
--endpoint https://gateway-adapter.example.com:8443

For an adapter that requires real routing identities, opt in with --delivery-fixture /path/to/delivery-fixture.json. The file must contain exactly one JSON object, at most 256 KiB, with only these fields (threadId is optional):

{
"accountId": "<test-account>",
"contextId": "<test-context>",
"threadId": "<test-thread>",
"replyTarget": "<test-reply-target>",
"originatingEventId": "<test-originating-event>"
}

Populate the placeholders only from retained delivery routing values you are authorized to inspect and reuse, for an explicitly approved test destination. Do not fabricate identities or use this flag to bypass adapter authorization. If you cannot obtain the retained values through authorized access, stop rather than switching identities or relaxing access controls. This check sends real messages, with fixed text [Orka conformance check] No action required. Without interim capability it sends one final plus its byte-identical duplicate. With capability it sends two distinct interim messages before final, duplicates the first interim and final, and replays the first interim after final to check stable receipts: three visible sends on a conforming adapter. It also adds an oversized-interim rejection probe only when capable; a non-capable adapter receives no message probes at all. Use an originating event that has not yet received terminal delivery for the interim check. Each invocation uses fresh delivery/idempotency IDs; running again can produce more visible messages or be rejected for an already-closed event. Do not fabricate a replacement event identity to bypass that constraint. The authentication and oversized-text/body rejection probes use the same routing identities. They should not deliver messages on a conforming adapter, but a broken adapter may deliver them. There are no automatic retries.

The fixture cannot supply text, credentials, delivery/idempotency IDs, or metadata. Unknown fields, malformed/trailing JSON, oversized files, and invalid routing identities fail before network requests. Fixture identities are redacted from checker results, and loader diagnostics do not include file paths or contents. Keep the fixture private; it is not a support-bundle artifact. Continue to supply the bearer token through the configured environment variable and keep TLS verification enabled. --delivery-fixture cannot be combined with --reference-fixtures, which sends reference-adapter fault deliveries. Without the new flag, the CLI retains its synthetic routing and fixed IDs; non-mutating readiness probes are unchanged.

The Gateway Live E2E GitHub Actions workflow deploys the TLS reference adapter, a deterministic external AgentRuntime, and a test-only native worker image in Kind. It preserves private-CA trust, readiness, ingress deduplication, and external-v2 final-only/no-Job coverage. Native coverage creates real gateway Tasks and controller Jobs/Pods, authenticates origin, and executes the production ReplyInConversationTool with the reusable native client. It verifies budget-backed admission/replay, observes a delivered interim receipt while the Task is still Running, and only then releases final completion. A rolled adapter with --interim-delivery=false must produce the tool's typed unsupported rejection with no enqueue and still deliver final. Local start/release fences are driven through Kubernetes exec; no production authorization exception is added. CI requires all three specs, retaining the zero-spec guard. It does not run the conformance CLI and does not replace conformance testing against each adapter build before rollout.

That deterministic fixture does not prove model-driven worker registration or live ACP tool execution. Native worker unit/integration tests separately cover actual registration, advertisement, dispatch, and continued model work. ACP integration tests exercise the production tool through a signed broker capability, active prompt/session guard, and real SQLite, with no native Job; Kubernetes and broker credential resolution use test fixtures, not a live RuntimePool/provider. Local unit/integration results and tagged E2E compilation are not evidence that the live suite ran. Live native delivery requires the workflow's three specs to execute successfully; the external-v2 spec remains final-only.

If a future release introduces another wire version, controller and adapter release notes must define an explicit dual-version overlap. Do not infer compatibility from similar payloads or from the adapter's product version.

SQLite schema support and rollback​

The controller creates the complete current SQLite schema for an empty database. For an existing database, startup checks the layout before serving gateway work. Current-layout stores retain their records across restarts. Incompatible layouts stop startup without conversion or replacement of the database. See the database support policy.

Before upgrading, rehearse startup against a copy of production data and compare Session/event/delivery counts, task references, terminal states, and PRAGMA integrity_check before and after reopening. Keep the pre-upgrade snapshot until the new release has processed queued events and deliveries successfully.

A binary-only rollback requires the prior release to support the retained SQLite layout and Kubernetes resources. Otherwise:

  1. Pause ingress and stop all controller writers.
  2. Restore the pre-upgrade SQLite/PVC snapshot.
  3. Restore the matching prior Gateway CRDs and Kubernetes objects without deleting stable object UIDs out from under retained ledger rows.
  4. Deploy the prior controller and adapter versions.
  5. Validate readiness and ledger counts before resuming ingress.
A controller that starts is not a controller that is safe

Use a controller release that supports the restored database layout and Kubernetes resources. This release provides no SQLite schema conversion in either direction.

Cleanup​

The maintenance loop marks pending work expired at its deadline, releases Session reservations, removes terminal deliveries older than retention (interim rows and their event ordering evidence remain until the owning event can be reclaimed), and compacts eligible terminal events into small deduplication tombstones before deleting their full ledger rows. It prunes the corresponding Gateway transcript messages once their event and reply records have expired. Active work, queued events, unsettled replies, and newer history keep the Session from being reclaimed.

For Sessions with ACP turns, retention records a durable cleanup intent, retires the exact Session runtime, and archives the turn receipts before removing the Session. Controller restart or a storage failure leaves the intent available for retry. Completed Tasks, including Tasks already waiting for deletion, can then finish their normal cleanup. Shared runtime pools remain available. The Kubernetes garbage collector may remove its own foregroundDeletion finalizer from an already-deleting Task; it cannot use that permission to change the Task or remove Orka's finalizer.

Compacted events keep an ownership receipt for their exact Task UID. The receipt stays while that Task exists, including during deletion and after the event tombstone expires. Maintenance removes it only after a direct Kubernetes API read confirms that the exact Task is gone. Read or storage errors leave the receipt available for retry.

Each maintenance pass examines at most --gateway-batch-size Session cleanup candidates and the same number of Task cleanup receipts, with a default of 25 and a maximum of 100. Later passes continue from the previous position. Blocked entries are retried after the scan wraps, so they do not keep later entries waiting indefinitely.

After a conversation is reclaimed, a new event starts a fresh Session with a different physical name. This also applies when a binding uses an explicit Session name: the configured name identifies the conversation, while the admitted event records its current physical Session name. Events arriving during cleanup receive a retryable response without consuming their external event ID. Old Session names and archive receipts remain reserved so delayed Task cleanup can still verify them.

Event tombstones expire after one additional retention window, so duplicate suppression remains bounded without retaining full message text forever. Replays within that window return the original event identity and do not start another turn, including after a fresh Session has started or a Task cleanup receipt has been pruned. Increase retention when audit requirements exceed 30 days. Size the persistent volume for the retained event, delivery, Session, Task-result, and tombstone windows, plus outstanding Task cleanup receipts and the Session completions and turn archives that remain reserved beyond those windows.