GCP Cloud Run deployment (Service Extensions)¶
This guide covers deploying Kysira when your stack is:
Kysira integrates as a Service Extensions traffic extension on the ALB. The LB forwards each request to kysira-ext-proc via the Envoy ext_proc gRPC protocol, which can inspect and optionally block the request before it reaches your Cloud Run service.
Architecture¶
┌─────────────────────────────────────┐
│ Google Cloud Application LB │
Internet ──────────▶│ │──────▶ Your Cloud Run service
│ Service Extension (ext_proc gRPC) │
└──────────────┬──────────────────────┘
│ gRPC (Envoy ext_proc protocol)
▼
kysira-ext-proc (Cloud Run)
│ HTTP
▼
kysira-inference (Cloud Run)
The LB is the single integration point. One kysira-ext-proc instance covers all of your Cloud Run backends — you do not need a sidecar per service.
Baseline this guide assumes¶
This guide adds Kysira to an app you already run on GCP. Before you start you should have:
- A Cloud Run service (your app) in some
REGION. - A Regional External Application Load Balancer (
load-balancing-scheme=EXTERNAL_MANAGED) fronting it — this is what the traffic extension attaches to. (On a Global External Application LB, follow Global External Application LB instead.)
The steps below reference YOUR_FORWARDING_RULE — that's the forwarding rule of that LB. Find its name with:
gcloud compute forwarding-rules list \
--filter="loadBalancingScheme=EXTERNAL_MANAGED" \
--format="table(name, region, IPAddress)"
Don't have an LB yet? Minimal setup (also what the deployment test uses)
A throwaway app + regional external ALB, enough to attach Kysira to. This is the exact baseline our on-demand deployment test stands up around a hello-world stand-in app.
# 1. A stand-in app on Cloud Run
gcloud run deploy demo-app \
--image gcr.io/cloudrun/hello \
--region REGION --allow-unauthenticated
# 2. Serverless NEG → backend service → URL map → target proxy → forwarding rule
gcloud compute network-endpoint-groups create demo-app-neg \
--region=REGION --network-endpoint-type=SERVERLESS \
--cloud-run-service=demo-app
gcloud compute backend-services create demo-app-bs \
--region=REGION --protocol=HTTP --load-balancing-scheme=EXTERNAL_MANAGED
gcloud compute backend-services add-backend demo-app-bs \
--region=REGION --network-endpoint-group=demo-app-neg \
--network-endpoint-group-region=REGION
gcloud compute url-maps create demo-app-urlmap \
--region=REGION --default-service=demo-app-bs
gcloud compute target-http-proxies create demo-app-proxy \
--region=REGION --url-map=demo-app-urlmap
# 3. A regional external ALB needs a proxy-only subnet in its region + VPC
# (one per region/network — skip if you already have one). Pick a range
# that doesn't overlap your existing subnets.
gcloud compute networks subnets create proxy-only-subnet \
--region=REGION --network=default \
--purpose=REGIONAL_MANAGED_PROXY --role=ACTIVE \
--range=192.168.240.0/23
gcloud compute forwarding-rules create demo-app-fr \
--region=REGION --load-balancing-scheme=EXTERNAL_MANAGED \
--network=default \
--target-http-proxy=demo-app-proxy --target-http-proxy-region=REGION \
--ports=80
demo-app-fr is then your YOUR_FORWARDING_RULE.
What you'll deploy¶
Two Cloud Run services, both stateless and CPU-only:
| Service | Image | Role |
|---|---|---|
kysira-inference | Provided in the Kysira Admin Console | Scores requests using ML models + regex detectors |
kysira-ext-proc | Provided in the Kysira Admin Console | Receives gRPC callouts from the ALB, calls inference, returns verdict |
Prerequisite — give Cloud Run access to the Kysira images¶
The Kysira release images live in Kysira's private Artifact Registry. Cloud Run pulls images as your project's Cloud Run service agent (a Google-managed identity), so the one-time setup is to grant that identity read access to our registry. You send Kysira your project number; we add the binding. No keys, no mirroring.
Which method do I use? (Cloud Run vs Kubernetes)
| Your platform | Use | Why |
|---|---|---|
| Cloud Run | Project number grant (this page) | Cloud Run pulls as a Google service agent and can't present a key — so we grant that identity access instead. |
| Kubernetes | Pull credential | K8s presents an imagePullSecret at pull time, so a JSON key is exactly what it wants. |
| Cloud Run, restricted egress / want to pin | Mirror (alternative below) | Use a pull credential to copy images into your own registry first. |
One-line rule: Kubernetes carries a key; Cloud Run can't — so Kubernetes gets a pull credential and Cloud Run gets a grant.
1. Find your project number¶
This is the numeric id of the project where Cloud Run runs (not the human-readable project id). It's not secret.
2. Register it in the Admin Console¶
In the Kysira Admin Console open Image Access → Add access → Google Cloud Run, paste the project number, and submit. Kysira grants your Cloud Run service agent (service-PROJECT_NUMBER@serverless-robot-prod.iam.gserviceaccount.com) roles/artifactregistry.reader on the releases repo.
If your project has never used Cloud Run
The service agent is created the first time you enable/use the Cloud Run API in the project. If the pull fails right after granting, deploy once (or enable the Cloud Run API), then retry — the grant activates as soon as the agent exists.
3. Use the Kysira image URLs directly¶
The Admin Console shows the full image paths. Use them as-is in the deploy commands below, in place of IMAGE_URL_FROM_KYSIRA_ADMIN_CONSOLE:
kysira-inference→us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/kysira-inference:TAGkysira-ext-proc→us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/kysira-ext-proc:TAG
No mirroring, no in-project repo — Cloud Run pulls straight from Kysira's registry once the grant is in place.
Alternative: mirror into your own registry¶
Prefer the project grant above for almost all cases. Mirror instead when you need to pin the exact image independently of your Kysira access, or when your environment has restricted egress that can't reach Kysira's registry at deploy time. Mirroring uses a pull credential (a JSON key) to copy the images into a repo in your own project, then Cloud Run pulls in-project.
Mirror steps
1. Get a pull credential. In the Admin Console open Image Access → Add access → Kubernetes / Docker, name it, and copy the JSON key — it is shown only once. Save it as kysira-key.json. The username is always _json_key; the JSON key is the password.
2. Authenticate Docker:
3. Create a repo in your project (skip if you already have one):
gcloud artifacts repositories create kysira \
--repository-format=docker \
--location=REGION \
--description="Mirrored Kysira agent images"
gcloud auth configure-docker REGION-docker.pkg.dev --quiet
4. Mirror the images. Replace REGION, YOUR_PROJECT, and TAG:
SRC=us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases
DST=REGION-docker.pkg.dev/YOUR_PROJECT/kysira
for img in kysira-inference kysira-ext-proc; do
docker pull "$SRC/$img:TAG"
docker tag "$SRC/$img:TAG" "$DST/$img:TAG"
docker push "$DST/$img:TAG"
done
Then use REGION-docker.pkg.dev/YOUR_PROJECT/kysira/<image>:TAG in the deploy commands below. Same-project Cloud Run pulls from this repo with no extra IAM.
Deployment phases¶
Start in Phase 1 (observabilityMode) for every new environment. Promote to Phase 3 once you have signal and trust.
| Phase | LB behaviour | Kysira mode | Traffic risk |
|---|---|---|---|
| 1 — Shadow (async) | LB ignores the verdict; callout is fire-and-forget | KYSIRA_MODE=shadow | Zero — the LB never waits on Kysira |
| 2 — Inline shadow | LB waits on the verdict but Kysira always returns CONTINUE | KYSIRA_MODE=shadow | Low — fail-open protects traffic even if Kysira is unreachable |
| 3 — Inline active | LB acts on the verdict; Kysira blocks above threshold | KYSIRA_MODE=active | Managed — fail-open still active |
Why Phase 1, not Phase 2?
Phase 1 (observabilityMode: true) is structurally incapable of affecting traffic — the LB never waits on Kysira and ignores its response. Phase 2 (inline + failOpen: true) protects against a Kysira outage but not against a miscalibrated model causing false positives. Phase 1 lets you validate scoring quality before any enforcement path exists.
Regional External Application LB required for Phase 1
observabilityMode: true is only supported by regional LbTrafficExtension resources. This guide provisions the callout backend service and traffic extension as regional, which requires a Regional External Application LB (a regional forwarding rule with load-balancing-scheme=EXTERNAL_MANAGED).
If you are running a Global External Application LB, follow Global External Application LB instead of the phases below. It has its own commands, and it replaces Phase 1 with an inline observation period in Kysira's shadow mode.
Phase 1 — Shadow observation (observabilityMode)¶
1. Deploy kysira-inference¶
gcloud run deploy kysira-inference \
--image IMAGE_URL_FROM_KYSIRA_ADMIN_CONSOLE \
--region REGION \
--port 8081 \
--cpu 2 \
--memory 4Gi \
--concurrency 12 \
--min-instances 1 \
--ingress internal \
--allow-unauthenticated \
--set-env-vars KYSIRA_DEVICE=cpu,KYSIRA_MAX_CONCURRENCY=4,OMP_NUM_THREADS=1
Inference needs at least 2 GB RAM (models are ~1.5 GB resident); 4 GiB leaves headroom for in-flight requests. Set --min-instances 1 to avoid cold-start latency on the scoring path. Scale-out takes tens of seconds, so size --min-instances for your normal peak if traffic is spiky — see Sizing inference. The --concurrency / KYSIRA_MAX_CONCURRENCY pair replaces Cloud Run's default of 80 concurrent requests per instance, which is far more than a model container can serve — see Sizing inference.
--ingress internal means Cloud Run rejects all requests from the public internet — only traffic that arrives through VPC is accepted. --allow-unauthenticated disables the IAM invoker check; the VPC ingress boundary is the security gate, so no OIDC token is required from ext-proc.
Licensing (protected models)¶
If you're running a licensed model, get a license key and an agent certificate from app.kysira.ai — see Licensing.
Store the credentials in Secret Manager:
cat client-cert.pem chain.pem > client-cert-chain.pem
echo -n "<your license token>" | gcloud secrets create kysira-license-token --data-file=-
gcloud secrets create kysira-license-cert --data-file=client-cert-chain.pem
gcloud secrets create kysira-license-key --data-file=client-key.pem
Then mount them and set the licensing env vars on the inference service:
gcloud run deploy kysira-inference \
--image IMAGE_URL_FROM_KYSIRA_ADMIN_CONSOLE \
--region REGION \
--port 8081 \
--cpu 2 \
--memory 4Gi \
--concurrency 12 \
--min-instances 1 \
--ingress internal \
--allow-unauthenticated \
--set-env-vars "KYSIRA_DEVICE=cpu,KYSIRA_MAX_CONCURRENCY=4,OMP_NUM_THREADS=1,LICENSE_ENFORCEMENT=permissive,AGENTS_BASE_URL=https://agents.kysira.ai,LICENSE_MODEL_ID=kysira/Argus-1,LICENSE_TOKEN_PATH=/etc/kysira/license-token/token,LICENSE_CLIENT_CERT_PATH=/etc/kysira/license-cert/client-cert.pem,LICENSE_CLIENT_KEY_PATH=/etc/kysira/license-key/client-key.pem" \
--set-secrets "/etc/kysira/license-token/token=kysira-license-token:latest,/etc/kysira/license-cert/client-cert.pem=kysira-license-cert:latest,/etc/kysira/license-key/client-key.pem=kysira-license-key:latest"
Each secret gets its own directory because Cloud Run mounts every secret as a separate volume, and two volumes can't share a mount directory. The LICENSE_*_PATH variables point the agent at those locations.
Start with LICENSE_ENFORCEMENT=permissive — the agent checks the license and logs the result, but boots either way. Once you've confirmed clean logs, redeploy with LICENSE_ENFORCEMENT=enforce so a denied license actually blocks the container from serving.
2. Deploy kysira-ext-proc¶
# Get the URL of the inference service
INFERENCE_URL=$(gcloud run services describe kysira-inference \
--region REGION --format 'value(status.url)')
gcloud run deploy kysira-ext-proc \
--image IMAGE_URL_FROM_KYSIRA_ADMIN_CONSOLE \
--region REGION \
--port 50051 \
--use-http2 \
--cpu 1 \
--memory 512Mi \
--min-instances 1 \
--allow-unauthenticated \
--vpc-egress all-traffic \
--network=YOUR_VPC_NETWORK \
--subnet=YOUR_VPC_SUBNET \
--set-env-vars "INFERENCE_URL=${INFERENCE_URL},KYSIRA_MODE=shadow,KYSIRA_SCORE_THRESHOLD=0.95"
--use-http2 is required — the ALB calls ext-proc over gRPC which requires HTTP/2.
--allow-unauthenticated is required. The Service Extensions callout from the load balancer does not carry a Cloud Run invoker (OIDC) token, so with --no-allow-unauthenticated Cloud Run rejects every callout (401 Empty Authorization header) — and because the traffic extension is failOpen, Kysira then silently stops enforcing while your traffic flows normally. The security boundary is that the callout backend service is only reachable via the load balancer, not an invoker check (same model as inference above).
--vpc-egress all-traffic routes all outbound traffic from ext-proc through your VPC. This is required for ext-proc to reach inference: *.run.app URLs resolve to public Google IPs, so with the default private-ranges-only egress those calls bypass the VPC and Cloud Run's ingress check on inference rejects them. With all-traffic, the calls go through VPC and are treated as internal.
Replace YOUR_VPC_NETWORK and YOUR_VPC_SUBNET with your VPC network name and subnet name (e.g. default and default). The subnet must be in the same region as the Cloud Run services.
The egress subnet must have Private Google Access enabled
With --vpc-egress all-traffic, ext-proc's calls to inference's *.run.app endpoint are routed through this subnet. If the subnet does not have Private Google Access enabled, that traffic has no private path to the endpoint and every call times out — ext-proc's circuit breaker opens and, because the traffic extension is failOpen, requests pass through unscored: clean traffic looks fine but nothing is actually being inspected. The default auto-mode subnet has PGA off by default. Enable it: gcloud compute networks subnets update YOUR_VPC_SUBNET --region REGION --enable-private-ip-google-access. Confirm ext-proc can reach inference — its logs should show scored decisions, not inference error / circuit open — requests passing through unscored.
3. Create the backend service for ext-proc¶
Service Extensions requires the callout target to be a backend service backed by a serverless NEG.
# Serverless NEG pointing to the ext-proc Cloud Run service
gcloud compute network-endpoint-groups create kysira-ext-proc-neg \
--region=REGION \
--network-endpoint-type=SERVERLESS \
--cloud-run-service=kysira-ext-proc
# Regional backend service (HTTP/2 for gRPC)
# Must be regional — observabilityMode requires a regional traffic extension
gcloud compute backend-services create kysira-ext-proc-bs \
--region=REGION \
--protocol=HTTP2 \
--load-balancing-scheme=EXTERNAL_MANAGED
gcloud compute backend-services add-backend kysira-ext-proc-bs \
--region=REGION \
--network-endpoint-group=kysira-ext-proc-neg \
--network-endpoint-group-region=REGION
4. Configure the traffic extension¶
Save this as traffic-extension-phase1.yaml:
name: projects/PROJECT_ID/locations/REGION/lbTrafficExtensions/kysira-waf
loadBalancingScheme: EXTERNAL_MANAGED
forwardingRules:
- projects/PROJECT_ID/regions/REGION/forwardingRules/YOUR_FORWARDING_RULE
extensionChains:
- name: kysira-chain
matchCondition:
celExpression: "true" # inspect all requests
extensions:
- name: kysira-ext-proc
authority: kysira-ext-proc
service: projects/PROJECT_ID/regions/REGION/backendServices/kysira-ext-proc-bs
timeout: 2s
failOpen: true # a Kysira outage never blocks your traffic
observabilityMode: true # Phase 1: async — LB ignores the verdict
supportedEvents:
- REQUEST_HEADERS
- REQUEST_BODY
Apply it:
gcloud service-extensions lb-traffic-extensions import kysira-waf \
--source=traffic-extension-phase1.yaml \
--location=REGION
observabilityMode is what makes Phase 1 free — not Kysira's shadow mode
Two different settings get called "shadow", and they do unrelated things:
| Setting | Owned by | What it controls |
|---|---|---|
KYSIRA_MODE=shadow | Kysira | Whether a flagged request gets a 403 or is only logged. No effect on latency. |
observabilityMode: true | Google Cloud | Whether the GFE waits for our answer at all. This is the latency switch. |
Service Extensions are inline by default. With observabilityMode omitted, the load balancer holds every request until the callout returns — whatever Kysira's own mode is set to, and regardless of failOpen. failOpen decides what happens after the wait, not whether there is one.
Keep observabilityMode: true for the whole validation window. It is the only configuration in which Kysira cannot affect your p95.
5. Verify¶
Check that Kysira is receiving traffic and scoring requests without affecting the LB:
# Stream ext-proc logs
gcloud run services logs read kysira-ext-proc --region REGION --tail 50
# Check inference health (no auth token needed — ingress is the boundary)
curl "${INFERENCE_URL}/health"
ext-proc logs structured JSON. You should see entries with "message": "decision" carrying score, action and mode fields. If you instead see inference error or inference circuit open — requests passing through unscored, ext-proc can't reach inference — check the Private Google Access warning above. No traffic impact to verify — in observabilityMode the LB is indifferent to Kysira's presence.
Inference is not reachable from your local machine
curl from your laptop will return HTTP 404 — Cloud Run's ingress is blocking the public request. This is correct behaviour. The health check above only succeeds from a VM in the same VPC network. Cloud Shell does not count: it runs outside your VPC, so it gets the same 404. The decision log entries are the easier proof that the path works end to end.
Graduating to live enforcement (Phase 3)¶
Once you've validated scoring quality in Phase 1 (typically 1–2 weeks of traffic), move to inline enforcement.
1. Update the traffic extension to inline mode¶
Save as traffic-extension-phase3.yaml:
name: projects/PROJECT_ID/locations/REGION/lbTrafficExtensions/kysira-waf
loadBalancingScheme: EXTERNAL_MANAGED
forwardingRules:
- projects/PROJECT_ID/regions/REGION/forwardingRules/YOUR_FORWARDING_RULE
extensionChains:
- name: kysira-chain
matchCondition:
celExpression: "true"
extensions:
- name: kysira-ext-proc
authority: kysira-ext-proc
service: projects/PROJECT_ID/regions/REGION/backendServices/kysira-ext-proc-bs
timeout: 0.5s # tune based on your p99 inference latency
failOpen: true # a Kysira outage → pass through, not block
# observabilityMode omitted (defaults false) — inline enforcement
supportedEvents:
- REQUEST_HEADERS
- REQUEST_BODY
gcloud service-extensions lb-traffic-extensions import kysira-waf \
--source=traffic-extension-phase3.yaml \
--location=REGION
The LB now waits on the Kysira verdict. failOpen: true means a Kysira outage or timeout causes the LB to pass the request through — not block it. This is the right default.
2. Kysira starts in shadow — no immediate enforcement¶
kysira-ext-proc defaults to KYSIRA_MODE=shadow, so it returns CONTINUE for every request while logging would-have-blocked decisions. Moving the extension inline does not change that — it only changes who waits.
Note that shadow mode does not rehearse active-mode latency. By default (KYSIRA_SHADOW_ASYNC=true) a shadow request is answered before it is scored, so an inline shadow deployment tells you nothing about what active mode will cost. To measure that, either:
- set
KYSIRA_SHADOW_ASYNC=false(withINFERENCE_TIMEOUT_MS=300, as in step 3 below) on a canary, which scores inline and gives you real inline timings while still never blocking; or - restrict the extension's
matchConditionto one low-traffic host and run it inline there first.
Verify false-positive rate from the decision logs (which are identical either way) and inline latency from the canary before switching KYSIRA_MODE=active.
ext-proc's HTTP endpoints (/_kysira/health, /api/mode, /metrics) listen on METRICS_PORT (9090), but Cloud Run only routes to the gRPC --port (50051), so they aren't reachable through the service URL. Confirm the mode from the config line each revision logs at startup instead:
gcloud logging read \
'resource.type=cloud_run_revision AND resource.labels.service_name=kysira-ext-proc AND jsonPayload.message="config"' \
--limit 1 --format 'value(resource.labels.revision_name, jsonPayload.mode)'
# → kysira-ext-proc-00002-abc shadow
3. Activate blocking¶
Mode is fixed at startup from KYSIRA_MODE and is not switchable at runtime — /api/mode is read-only (GET; a POST returns 405). Flipping it means a new revision:
gcloud run services update kysira-ext-proc \
--region REGION \
--update-env-vars KYSIRA_MODE=active,INFERENCE_TIMEOUT_MS=300,KYSIRA_EXTENSION_TIMEOUT_MS=500
Inline scoring needs INFERENCE_TIMEOUT_MS=300 (the default from v0.7.1)
v0.7.0 defaulted inline to 100 ms, which leaves the sidecar an 85 ms budget per request. On the 2 vCPU sizing in this guide an Argus pass takes longer than that under load, so the sidecar answers almost every request regex-only (degraded: true) — blocking still happens, but the ML model is effectively off. v0.7.1 defaults inline to 300 ms; setting it explicitly, as above, is harmless on v0.7.1 and required on v0.7.0. 300 ms fits a model pass and stays below the traffic extension's 0.5 s timeout (see Timeout ordering). KYSIRA_EXTENSION_TIMEOUT_MS makes ext-proc check that ordering at startup. After activating, watch the share of degraded decisions — see Sizing inference.
Confirm the new revision is actually serving before trusting the flip — the newest config line should name the new revision with active, and that revision should hold the traffic:
gcloud logging read \
'resource.type=cloud_run_revision AND resource.labels.service_name=kysira-ext-proc AND jsonPayload.message="config"' \
--limit 1 --format 'value(resource.labels.revision_name, jsonPayload.mode)'
# → kysira-ext-proc-00003-xyz active
gcloud run services describe kysira-ext-proc --region REGION \
--format 'value(status.traffic)'
Rollback — decide this before you activate, because there is no instant revert. The fastest path is a traffic shift back to the known-good revision: it takes effect in seconds and needs no new build.
# list revisions, newest first
gcloud run revisions list --service kysira-ext-proc --region REGION --limit 5
# send all traffic back to the previous one
gcloud run services update-traffic kysira-ext-proc \
--region REGION --to-revisions PREVIOUS_REVISION=100
Setting KYSIRA_MODE=shadow with services update also works but rolls a new revision, so it is slower. Either way there is a window of tens of seconds in which blocking is still in effect. That window is exactly why activation should begin as a canary traffic split rather than a whole-service flip.
Global External Application LB¶
Use this section if your app sits behind a Global External Application Load Balancer. To check, list your global forwarding rules; your LB's rule appears here if it's global:
gcloud compute forwarding-rules list --global \
--filter='loadBalancingScheme=EXTERNAL_MANAGED' \
--format='table(name, IPAddress, target.basename())'
The Kysira services are the same as on a regional LB. Three things differ:
| Regional LB | Global LB | |
|---|---|---|
| Callout backend service | regional (--region=REGION) | global (--global) |
| Traffic extension | --location=REGION, regions/... paths | --location=global, global/... paths |
Phase 1 (observabilityMode) | available | not available: it only exists on regional traffic extensions |
Because there is no Phase 1, a Global LB starts inline with Kysira in shadow mode. The LB waits on Kysira for every request, but Kysira always answers CONTINUE and only logs what it would have blocked. failOpen: true passes traffic through if Kysira is down or slow. This is the same as Phase 2 in Deployment phases: no false positive can block traffic, but the extension's latency is in the request path from day one. Keep the extension timeout short and watch your p95 while you observe.
A global LB also needs no proxy-only subnet. That requirement only applies to regional LBs.
1. Deploy kysira-inference and kysira-ext-proc¶
Follow Phase 1, step 1 and step 2 unchanged. Cloud Run services are always regional, so pick the REGION closest to most of your traffic (see Regional co-location). Leave KYSIRA_MODE at its default, shadow.
2. Create a global backend service for ext-proc¶
The serverless NEG is regional, as it always is. The backend service that the traffic extension calls must be global to match the LB:
gcloud compute network-endpoint-groups create kysira-ext-proc-neg \
--region=REGION \
--network-endpoint-type=SERVERLESS \
--cloud-run-service=kysira-ext-proc
# Global backend service (HTTP/2 for gRPC). --port-name=http is required:
# for a global backend service gcloud otherwise derives the port name from the
# protocol ("http2"), and serverless NEGs reject any port name but the default.
gcloud compute backend-services create kysira-ext-proc-bs \
--global \
--protocol=HTTP2 \
--port-name=http \
--load-balancing-scheme=EXTERNAL_MANAGED
gcloud compute backend-services add-backend kysira-ext-proc-bs \
--global \
--network-endpoint-group=kysira-ext-proc-neg \
--network-endpoint-group-region=REGION
3. Attach the traffic extension (inline)¶
Save as traffic-extension-global.yaml:
name: projects/PROJECT_ID/locations/global/lbTrafficExtensions/kysira-waf
loadBalancingScheme: EXTERNAL_MANAGED
forwardingRules:
- projects/PROJECT_ID/global/forwardingRules/YOUR_FORWARDING_RULE
extensionChains:
- name: kysira-chain
matchCondition:
celExpression: "true" # inspect all requests
extensions:
- name: kysira-ext-proc
authority: kysira-ext-proc
service: projects/PROJECT_ID/global/backendServices/kysira-ext-proc-bs
timeout: 0.5s # the LB waits at most this long per request
failOpen: true # a Kysira outage or timeout passes traffic through
# no observabilityMode: global extensions are always inline
supportedEvents:
- REQUEST_HEADERS
- REQUEST_BODY
gcloud service-extensions lb-traffic-extensions import kysira-waf \
--source=traffic-extension-global.yaml \
--location=global
The extension can take a few minutes to reach every Google Front End.
4. Verify and observe in shadow¶
Check that ext-proc logs decision entries, as in Phase 1, step 5. The logs are on the Cloud Run service in REGION, whichever edge location served the request.
Stay in shadow for the same 1–2 weeks you would spend in Phase 1. Review the would-have-blocked decisions for false positives. To measure what active mode will add to latency before you turn it on, run a canary with KYSIRA_SHADOW_ASYNC=false, as described in Kysira starts in shadow.
5. Activate blocking¶
Identical to Phase 3, step 3: set KYSIRA_MODE=active on ext-proc, confirm the new revision is serving, and have the rollback command ready.
Serving several regions¶
A global LB takes traffic worldwide, but the steps above run Kysira in one region, so requests from far-away users make a longer round trip to score. To keep scoring close to your users, deploy inference and ext-proc in each region you serve, then add one serverless NEG per region to the same global backend service:
gcloud compute network-endpoint-groups create kysira-ext-proc-neg-REGION2 \
--region=REGION2 \
--network-endpoint-type=SERVERLESS \
--cloud-run-service=kysira-ext-proc
gcloud compute backend-services add-backend kysira-ext-proc-bs \
--global \
--network-endpoint-group=kysira-ext-proc-neg-REGION2 \
--network-endpoint-group-region=REGION2
The LB sends each callout to the nearest healthy region. You still need only one traffic extension.
Environment variable reference¶
kysira-ext-proc¶
| Variable | Default | Description |
|---|---|---|
INFERENCE_URL | http://localhost:8081 | URL of the kysira-inference service |
KYSIRA_MODE | shadow | shadow (log only) or active (block above threshold) |
KYSIRA_SCORE_THRESHOLD | 0.95 | Kill threshold [0–1] |
KYSIRA_LOG_BODY | false | Set to true to include body_sample in every decision log, not just flagged requests. Useful during initial rollout to validate what the model sees. |
KYSIRA_LOG_HEADERS | true | Include the request headers on every decision log as a headers object. Values of sensitive-named headers (Authorization, Cookie, X-API-Key, …) are replaced with {redacted}; names are always kept. Set to false to omit the field. |
KYSIRA_REDACT_KEYS | — | Extra comma-separated field-name substrings whose query/body/header values are replaced with {redacted} in decision logs, on top of the built-in defaults (token, secret, key, email, …). Case-insensitive substring match. See observability.md. |
INFERENCE_TIMEOUT_MS | 500 async shadow / 300 inline | Milliseconds ext-proc will wait for an inference response before failing open. The default depends on placement: async shadow is off the critical path and takes 500 for coverage; inline (active mode, or KYSIRA_SHADOW_ASYNC=false) adds every millisecond to the caller, so it defaults to a 300 ms latency budget (100 ms in v0.7.0, which is too tight for Argus on the sizing in this guide — set 300 explicitly there; see Activate blocking). ext-proc forwards it to inference as X-Kysira-Budget-Ms, minus a 15 ms margin for the round trip that sits inside this deadline but outside the sidecar's clock. A call inference cannot finish inside that budget comes back regex-only (degraded: true, counted in kysira_extproc_inference_degraded_total) rather than queueing for a verdict that would arrive too late. Must be lower than the traffic extension's timeout — see Timeout ordering. |
KYSIRA_SHADOW_ASYNC | true | In shadow mode, answer the load balancer immediately and score in the background so scoring never adds latency. Set false to score inline — useful as a latency rehearsal before switching to active mode, at the cost of that latency. Ignored in active mode, which is always inline. |
KYSIRA_SHADOW_WORKERS | 32 | Concurrent background scorers when async shadow is on. Also bounds ext-proc's connections to inference. |
KYSIRA_SHADOW_QUEUE | 2048 | Background scoring queue depth. When full, jobs are dropped rather than delaying traffic — traffic is never affected, only scoring coverage. Drops increment kysira_extproc_shadow_dropped_total and emit a rate-limited WARNING carrying the cumulative count. |
KYSIRA_BREAKER_THRESHOLD | 5 | Consecutive inference failures before the circuit opens and requests skip scoring entirely instead of each waiting out the timeout. 0 disables. |
KYSIRA_BREAKER_COOLDOWN_MS | 10000 | How long the circuit stays open before admitting one probe request. |
KYSIRA_EXTENSION_TIMEOUT_MS | — | Advisory. Tell ext-proc the traffic extension timeout you configured and it will warn at startup if INFERENCE_TIMEOUT_MS is not safely below it. |
KYSIRA_MAX_BODY_BYTES | 1048576 | Cap on request body accumulated for scoring. Larger bodies are truncated for scoring only and pass through untouched. |
EXT_PROC_PORT | 50051 | gRPC listen port (set Cloud Run --port to match) |
METRICS_PORT | 9090 | HTTP port for /_kysira/health and /metrics |
kysira-inference¶
| Variable | Default | Description |
|---|---|---|
KYSIRA_DEVICE | auto | cpu, cuda, or mps — use cpu on Cloud Run |
KYSIRA_MAX_CONCURRENCY | 4 | Concurrent scoring requests allowed into the model. Size against the CPU limit and set Cloud Run's --concurrency to match — see Sizing inference. |
KYSIRA_ADMISSION | on | Kill switch. With off, /score/all queues on the threadpool exactly as it did before admission control existed — the rollback that does not need the previous image. |
KYSIRA_MAX_QUEUE | 2 × KYSIRA_MAX_CONCURRENCY | How many /score/all calls may wait for a model slot when none is free. Beyond this the call does not join the queue: the regex/signature detectors run anyway and the verdict comes back 200 with degraded: true and no ML heads. 0 means never queue — degrade as soon as all slots are busy (a free slot is always used, idle sidecars are never degraded). |
KYSIRA_QUEUE_TIMEOUT_MS | 250 | Longest a /score/all call waits for a slot. The effective wait is the smaller of this and the caller's X-Kysira-Budget-Ms minus the estimated scoring time; a call that cannot both wait and finish inside the budget is degraded rather than started, since a verdict delivered after the caller's deadline is wasted compute the caller records as a failure. |
KYSIRA_ARGUS_TRIM_MIN_RUN | 64 | Argus input trim: an opaque run (JWT segment, signature, session id) of at least this many characters in the appended Authorization/Cookie lines is cut to its first KYSIRA_ARGUS_TRIM_KEEP characters before tokenizing. Cuts a bearer-token request from ~400 tokens to ~140 with no benign FP change on the v4 eval set. 0 disables. Regex detectors always see the full value. |
KYSIRA_ARGUS_TRIM_KEEP | 32 | Characters kept from the start of each trimmed run. |
KYSIRA_TORCH_THREADS | 1 | Torch intra-op threads per request. Leave at 1 unless you have raised KYSIRA_MAX_CONCURRENCY well below the core count. |
KYSIRA_SCORE_THRESHOLD | (not used here) | Threshold is enforced by ext-proc, not inference |
KYSIRA_ARGUS_DISABLE | "" | Comma-separated Argus sub-models to suppress (e.g. generic_attack) |
LICENSE_ENFORCEMENT | off | off, permissive (verify + log, still boot), or enforce (fail closed). Any other value stops the container at startup, so a typo cannot silently turn the gate off. See Licensing. |
KYSIRA_MAX_REQUEST_TEXT | 131072 | Most characters of one request's text the detectors will look at. Longer text is clipped to its head and tail (never rejected, so an oversized request cannot trip the caller's circuit breaker). The detector battery is linear but not free, so this bounds the worst case one unauthenticated request can cost. 0 disables clipping. |
AGENTS_BASE_URL | — | Kysira key-delivery endpoint. Set to https://agents.kysira.ai |
LICENSE_MODEL_ID | kysira/Argus-1 | The model your license is entitled to |
LICENSE_TOKEN_PATH | /etc/kysira/license/token | Path to the mounted license token |
LICENSE_CLIENT_CERT_PATH | /etc/kysira/license/client-cert.pem | Path to the mounted certificate + CA chain |
LICENSE_CLIENT_KEY_PATH | /etc/kysira/license/client-key.pem | Path to the mounted certificate private key |
Regional co-location¶
Place kysira-ext-proc in the same region as your Cloud Run workloads. Since the traffic extension and backend service are regional, each region has its own independent stack. For multi-region deployments, repeat the Phase 1 and Phase 3 steps once per region — create a separate NEG, backend service, and traffic extension in each region, each pointing at that region's forwarding rule. On a Global LB, use one global backend service with a NEG per region instead; see Serving several regions.
Timeout ordering¶
Three timeouts sit inside one another on the request path. They only behave as intended if each inner one is strictly smaller than the one wrapping it:
traffic extension `timeout` 500ms ← the load balancer's patience
└─ INFERENCE_TIMEOUT_MS 300ms ← ext-proc's patience
└─ observed inference p99 ~80ms ← what the model actually takes
Invert any pair and the inner layer's fail-open becomes unreachable code: the outer layer gives up first, so instead of failing open quickly at the inner deadline, every slow request pays the outer timeout in full. An INFERENCE_TIMEOUT_MS of 3000 under a 2s extension timeout does not give inference more room — it just guarantees a 2-second wait whenever inference is slow.
Set KYSIRA_EXTENSION_TIMEOUT_MS to whatever you configured as the extension timeout and ext-proc will check this for you at startup, logging a WARNING if the ordering is wrong.
This ordering only matters when ext-proc is inline. Under observabilityMode: true, or in shadow mode with the default KYSIRA_SHADOW_ASYNC=true, nothing is waiting on the score.
Sizing inference¶
kysira-inference is CPU-bound and holds memory per in-flight request, so it is sized by concurrency, not by memory alone. Two defaults will hurt you if left alone:
- Cloud Run's request concurrency defaults to 80. That is how many requests one instance accepts before Cloud Run considers scaling out. A model container degrades long before 80, so the scale-out signal never fires — the instance just gets slower and then OOMs.
- anyio's threadpool defaults to 40. The scoring endpoints are sync handlers, so without a cap that is 40 concurrent forward passes on however many cores you gave the container.
Set both, and keep them in the same neighbourhood:
gcloud run deploy kysira-inference \
--region REGION \
--cpu 2 --memory 4Gi \
--concurrency 12 \
--min-instances 1 --max-instances 10 \
--set-env-vars KYSIRA_MAX_CONCURRENCY=4,OMP_NUM_THREADS=1
A reasonable starting point is KYSIRA_MAX_CONCURRENCY ≈ 2× vCPU and Cloud Run --concurrency = KYSIRA_MAX_CONCURRENCY + KYSIRA_MAX_QUEUE (3× with the default queue — 12 above), then load-test at your real peak QPS.
Scale-out is not instant — size --min-instances for your normal peak. A new inference instance takes tens of seconds to become ready (in our deployment test, 25–50 s: container start plus ~17 s of model load), and until it is, the instances already running carry all the traffic. A spike beyond what they can serve shows up as timeouts, circuit-breaker opens and requests passing unscored (or regex-only) for that window. Keep enough warm instances for your usual peak rather than relying on --min-instances 1 and autoscaling. Raising memory without capping concurrency moves the cliff; it does not remove it.
Inference will not let itself be the cliff: past KYSIRA_MAX_CONCURRENCY in-flight calls, up to KYSIRA_MAX_QUEUE more wait (at most KYSIRA_QUEUE_TIMEOUT_MS, or less if the caller's X-Kysira-Budget-Ms says it would be too late anyway). Everything beyond that is degraded, not dropped: the regex and signature detectors still run — they cost microseconds and need no model slot — and the verdict returns 200 with degraded: true. So a capacity shortfall costs the ML heads, not detection altogether, and never becomes latency on the request path.
This matters for more than tidiness. Answering 503 instead would have handed an attacker a way to switch the models off with a burst of load and then send a payload through the gap, and would have fed the caller's circuit breaker with what is really a healthy, fast answer.
Watch it as lost model coverage: kysira_inference_degraded_total{reason} on inference, kysira_extproc_inference_degraded_total and the degraded field on decision log lines on ext-proc. Every decision line also carries the sidecar's own split of the time (sidecar_queue_ms, sidecar_score_ms, tokens): if sidecar_queue_ms is where the latency is, inference needs capacity, not a faster model.
Watch kysira_extproc_request_duration_seconds for the tail, kysira_extproc_inference_degraded_total for capacity, and kysira_extproc_breaker_opens_total for inference falling over.
Set Cloud Run --concurrency to at least KYSIRA_MAX_CONCURRENCY + KYSIRA_MAX_QUEUE. Below that, Cloud Run queues the overflow outside the sidecar, where it is invisible to admission control and to the budget — the shaping described above never gets to run.
Fail-open security posture¶
failOpen: true means a Kysira outage, timeout, or error causes the ALB to pass the request through rather than block it. This is the correct default — it prevents Kysira from becoming a single point of failure for your production traffic.
The tradeoff: under a sustained Kysira outage, enforcement silently disappears. Mitigate with:
- Cloud Run
--min-instances 1to eliminate cold starts on the scoring path - Alerting on
kysira_extproc_inference_errors_totalandkysira_extproc_breaker_opens_total(exposed atMETRICS_PORT/metrics). On Cloud Run that port isn't reachable from outside the container — scrape it with a Managed Service for Prometheus sidecar, or alert with log-based metrics on theinference errorandinference circuit openedlog messages timeouttuned conservatively relative to your observed inference p99, and ordered correctly againstINFERENCE_TIMEOUT_MS(see Timeout ordering)
If your threat model requires fail-secure behaviour in active mode, set failOpen: false — but only after validating that Kysira's availability SLA meets your traffic SLA.
Common errors¶
gcloud run deploy fails with "invalid image" or "image not found"
The --image IMAGE_URL_FROM_KYSIRA_ADMIN_CONSOLE placeholder must be replaced with a real image URL — the Kysira releases path (us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/kysira-inference:TAG) if you used a project grant, or your mirrored URL (REGION-docker.pkg.dev/YOUR_PROJECT/kysira/kysira-inference:TAG) if you mirrored. See Prerequisite — give Cloud Run access to the Kysira images. The deploy fails immediately if the literal placeholder string is used.
gcloud run deploy fails with "permission denied" pulling the image
Cloud Run pulls as your project's service agent (service-PROJECT_NUMBER@serverless-robot-prod.iam.gserviceaccount.com). If you used a project grant, confirm you sent the numeric project number of the project where Cloud Run runs, and that the project has used Cloud Run at least once (the service agent is created lazily). If you mirrored into a repo in a different project than the Cloud Run service, that project's service agent needs roles/artifactregistry.reader on the repo — mirroring into the same project avoids this.