Troubleshooting¶
Inference pod not becoming ready¶
Symptom: inference pod stays in Init or readiness probe fails¶
The inference container loads two ML models (~1.5 GB combined) on startup. This takes 2–5 minutes on a cold pull and 30–60 seconds on a warm node with cached layers. The readiness probe has a 3-minute initial delay to account for this.
# Watch pod status
kubectl get pods -n kysira -w
# Stream inference logs to see load progress
kubectl logs -n kysira -l app.kubernetes.io/name=kysira-inference -f
Look for Model loaded or status: ok in the logs. If you see a Python traceback instead:
-
OOMKilled — the node doesn't have enough memory. The two models need ~1.5 GB resident. Increase the memory limit:
Then increase via Helm: -
Image pull error — the
kysira-pullsecret may be missing or expired. Re-create it with a fresh JSON key from app.kysira.ai:
High inference latency / scoring timeouts¶
Symptom: ext-proc logs inference error: context deadline exceeded¶
Two different timeouts govern a scored request, and this error is the inner one firing:
| Timeout | Set by | Bounds |
|---|---|---|
envoyFilter.messageTimeout (500ms) | Envoy | How long Envoy waits for ext-proc to answer |
config.inferenceTimeoutMs (300) | ext-proc | How long ext-proc waits for a score |
The message above means the inner one elapsed: ext-proc gave up on inference and failed the request open. If inference is slower than that:
# Check which device inference is using
kubectl exec -n kysira deployment/kysira-inference -- \
wget -qO- http://localhost:8081/health | python3 -m json.tool
On CPU, the prompt-injection classifier takes 200–400 ms per request. Options:
-
GPU — set
KYSIRA_DEVICE=cudaand add a GPU resource limit: -
DaemonSet mode — one inference pod per node eliminates the network hop:
bash helm upgrade kysira oci://us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/charts/kysira-platform \ --namespace kysira --reuse-values \ --set "kysira-inference.daemonSet.enabled=true" -
Cap concurrency — this is usually the real fix, and raising the timeout only hides it. Inference is CPU-bound and its latency degrades sharply once more requests are in flight than the container has cores for. Set
KYSIRA_MAX_CONCURRENCY(default4) against the CPU limit, and set the serving platform's own request concurrency to match — on Cloud Run,--concurrency, whose default of 80 is far too high for a model container. See Sizing inference. -
Increase timeout — raise
config.inferenceTimeoutMs(chart default300; the binary falls back to500when the env var is unset). Do this only if timeouts are occasional rather than systematic, and keep it strictly belowenvoyFilter.messageTimeout, or the increase does nothing but lengthen the wait — Envoy gives up first and ext-proc's fail-open never runs. See Timeout ordering.
ext-proc fails open on timeout — the request passes through with an error logged, never dropped.
Symptom: fewer decisions logged than requests served¶
Two things silently reduce scoring coverage, and both now say so in the logs. Check for the circuit breaker first (inference circuit open, above) — under load it is far the more likely of the two, since it sheds every request for a full cooldown once inference starts failing.
The other is the shadow queue. In shadow mode a request that cannot be queued for background scoring is passed through unscored. This never affects traffic — the response was already sent — but it does mean the decision log undercounts. Look for:
It carries dropped_total, and is rate-limited to one line per 10s so a sustained overload does not become a logging flood. Sustained drops mean inference cannot keep up with request volume: scale inference (see Sizing inference) or raise KYSIRA_SHADOW_WORKERS. Raising KYSIRA_SHADOW_QUEUE alone only buys a longer queue in front of the same bottleneck.
If inference is failing rather than merely slow, ext-proc's circuit breaker opens after KYSIRA_BREAKER_THRESHOLD consecutive failures (default 5) and requests skip scoring entirely instead of each waiting out the timeout. This is logged, so you do not need /metrics to see it:
WARNING inference circuit opened — skipping inference until it recovers
WARNING inference circuit open — requests passing through unscored (shed_total=…)
INFO inference circuit closed — scoring resumed (shed_total=…)
The middle line repeats at most once per 10s while the circuit is open and carries a running shed_total. Requests shed this way are passed through unscored — traffic is unaffected, but they produce no decision log entry, so the decision log will undercount for as long as the circuit is open. That is the usual explanation for a gap between requests served and decisions logged.
Slow scoring should not be slow traffic
In shadow mode ext-proc answers before it scores (KYSIRA_SHADOW_ASYNC=true, the default), so inference latency stays off the request path entirely. If you are seeing client-visible latency from scoring in shadow mode, check that setting first.
On a GCP load balancer, a regional traffic extension can also set observabilityMode: true, which stops the GFE waiting on the callout at all. It is not available on a global external Application Load Balancer, where the callout is always inline — there, async shadow is the only thing keeping scoring off the request path.
False positives (legitimate requests flagged)¶
Symptom: normal API calls appear in the dashboard with high scores¶
# Check what the inference service returns for a specific payload
kubectl exec -n kysira deployment/kysira-inference -- \
wget -qO- --post-data='{"request_text":"your payload here"}' \
--header='Content-Type: application/json' \
http://localhost:8081/score/all | python3 -m json.tool
The detector field in the response tells you which classifier fired. Common causes:
-
sqli— the SQL injection classifier can be over-eager on SQL keywords in prose. Raise the threshold:helm upgrade kysira oci://us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/charts/kysira-platform \ --namespace kysira --reuse-values \ --set "kysira-proxy.config.scoreThreshold=0.98"0.98significantly reduces false positives at some cost to recall. -
xss— the regex fires on<script,javascript:,onerror=, and similar. If your app legitimately POSTs HTML content, route those specific paths outside Kysira. -
nosqli— legitimate Mongo$in/$gte/$ltefilter operators in JSON bodies trigger an advisory score (below the kill line by default). They only block in active mode if you've lowered the threshold below0.7.
Active mode not blocking requests¶
Symptom: mode is active but malicious requests still reach your app¶
-
Confirm the mode change took effect:
-
Check the dashboard — if
actionshowsshadow_killinstead ofactive_kill, the pod is still in shadow mode. The mode API sets mode in-memory per pod; use Helm to persist it across restarts:bash helm upgrade kysira oci://us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/charts/kysira-platform \ --namespace kysira --reuse-values \ --set "kysira-proxy.config.mode=active" -
If
actionshowspassed, the score is below the threshold — the request is genuinely not being flagged. Lower the threshold or test with a more obvious payload like' OR 1=1--. -
ext-proc only — the mode API is on the HTTP port (
:9090), not the gRPC port, and it is read-only: unlike the proxy, ext-proc fixes its mode at startup fromKYSIRA_MODEand has no write path, so aPOSTreturns405. Changing the mode means redeploying with the new value. Confirm what the running pod is actually set to:
Dashboard shows no events¶
Symptom: dashboard loads but the event feed is empty¶
# Check the proxy is healthy and the dashboard can reach it
kubectl logs -n kysira -l app.kubernetes.io/name=kysira-dashboard
The dashboard proxies /api/ requests to the kysira-proxy service. Upstream connection errors here mean the proxy service name or port is wrong in the dashboard ConfigMap — check the proxyServiceName Helm value matches the proxy service name:
If the service names look right, check the proxy itself is receiving traffic:
HPA not scaling inference¶
Symptom: inference pod count stays at 1 under load¶
The HPA requires the Metrics Server. Check if it's installed:
If this fails, install the Metrics Server:
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
Then check the HPA status:
The HPA targets 80% CPU by default. If inference is GPU-bound, CPU utilization will be low and the HPA won't trigger — set a fixed replica count or use a custom Prometheus-based metric instead.
Checking metrics directly¶
# Proxy metrics
kubectl exec -n kysira deployment/kysira-proxy -- \
wget -qO- http://localhost:8080/metrics | grep kysira_
# ext-proc metrics
kubectl exec -n kysira deployment/kysira-ext-proc -- \
wget -qO- http://localhost:9090/metrics | grep kysira_extproc_
# Inference metrics
kubectl exec -n kysira deployment/kysira-inference -- \
wget -qO- http://localhost:8081/metrics | grep kysira_inference_
See Observability for connecting these to Grafana or Datadog.