Skip to content

Troubleshooting


Inference pod not becoming ready

Symptom: inference pod stays in Init or readiness probe fails

The inference container loads two ML models (~1.5 GB combined) on startup. This takes 2–5 minutes on a cold pull and 30–60 seconds on a warm node with cached layers. The readiness probe has a 3-minute initial delay to account for this.

# Watch pod status
kubectl get pods -n kysira -w

# Stream inference logs to see load progress
kubectl logs -n kysira -l app.kubernetes.io/name=kysira-inference -f

Look for Model loaded or status: ok in the logs. If you see a Python traceback instead:

  • OOMKilled — the node doesn't have enough memory. The two models need ~1.5 GB resident. Increase the memory limit:

    kubectl describe pod -n kysira -l app.kubernetes.io/name=kysira-inference | grep -A5 OOMKilled
    
    Then increase via Helm:
    helm upgrade kysira oci://us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/charts/kysira-platform \
      --namespace kysira --reuse-values \
      --set "kysira-inference.resources.limits.memory=4Gi" \
      --set "kysira-inference.resources.requests.memory=2Gi"
    

  • Image pull error — the kysira-pull secret may be missing or expired. Re-create it with a fresh JSON key from app.kysira.ai:

    kubectl delete secret kysira-pull -n kysira
    kubectl create secret docker-registry kysira-pull \
      --docker-server=us-central1-docker.pkg.dev \
      --docker-username=_json_key \
      --docker-password='<paste your JSON key>' \
      --namespace kysira
    


High inference latency / scoring timeouts

Symptom: ext-proc logs inference error: context deadline exceeded

Two different timeouts govern a scored request, and this error is the inner one firing:

Timeout Set by Bounds
envoyFilter.messageTimeout (500ms) Envoy How long Envoy waits for ext-proc to answer
config.inferenceTimeoutMs (300) ext-proc How long ext-proc waits for a score

The message above means the inner one elapsed: ext-proc gave up on inference and failed the request open. If inference is slower than that:

# Check which device inference is using
kubectl exec -n kysira deployment/kysira-inference -- \
  wget -qO- http://localhost:8081/health | python3 -m json.tool

On CPU, the prompt-injection classifier takes 200–400 ms per request. Options:

  1. GPU — set KYSIRA_DEVICE=cuda and add a GPU resource limit:

    kysira-inference:
      config:
        device: cuda
      resources:
        limits:
          nvidia.com/gpu: "1"
    

  2. DaemonSet mode — one inference pod per node eliminates the network hop: bash helm upgrade kysira oci://us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/charts/kysira-platform \ --namespace kysira --reuse-values \ --set "kysira-inference.daemonSet.enabled=true"

  3. Cap concurrency — this is usually the real fix, and raising the timeout only hides it. Inference is CPU-bound and its latency degrades sharply once more requests are in flight than the container has cores for. Set KYSIRA_MAX_CONCURRENCY (default 4) against the CPU limit, and set the serving platform's own request concurrency to match — on Cloud Run, --concurrency, whose default of 80 is far too high for a model container. See Sizing inference.

  4. Increase timeout — raise config.inferenceTimeoutMs (chart default 300; the binary falls back to 500 when the env var is unset). Do this only if timeouts are occasional rather than systematic, and keep it strictly below envoyFilter.messageTimeout, or the increase does nothing but lengthen the wait — Envoy gives up first and ext-proc's fail-open never runs. See Timeout ordering.

ext-proc fails open on timeout — the request passes through with an error logged, never dropped.

Symptom: fewer decisions logged than requests served

Two things silently reduce scoring coverage, and both now say so in the logs. Check for the circuit breaker first (inference circuit open, above) — under load it is far the more likely of the two, since it sheds every request for a full cooldown once inference starts failing.

The other is the shadow queue. In shadow mode a request that cannot be queued for background scoring is passed through unscored. This never affects traffic — the response was already sent — but it does mean the decision log undercounts. Look for:

"shadow scoring queue full — requests passed through unscored"

It carries dropped_total, and is rate-limited to one line per 10s so a sustained overload does not become a logging flood. Sustained drops mean inference cannot keep up with request volume: scale inference (see Sizing inference) or raise KYSIRA_SHADOW_WORKERS. Raising KYSIRA_SHADOW_QUEUE alone only buys a longer queue in front of the same bottleneck.


If inference is failing rather than merely slow, ext-proc's circuit breaker opens after KYSIRA_BREAKER_THRESHOLD consecutive failures (default 5) and requests skip scoring entirely instead of each waiting out the timeout. This is logged, so you do not need /metrics to see it:

WARNING  inference circuit opened — skipping inference until it recovers
WARNING  inference circuit open — requests passing through unscored   (shed_total=…)
INFO     inference circuit closed — scoring resumed                   (shed_total=…)

The middle line repeats at most once per 10s while the circuit is open and carries a running shed_total. Requests shed this way are passed through unscored — traffic is unaffected, but they produce no decision log entry, so the decision log will undercount for as long as the circuit is open. That is the usual explanation for a gap between requests served and decisions logged.

Slow scoring should not be slow traffic

In shadow mode ext-proc answers before it scores (KYSIRA_SHADOW_ASYNC=true, the default), so inference latency stays off the request path entirely. If you are seeing client-visible latency from scoring in shadow mode, check that setting first.

On a GCP load balancer, a regional traffic extension can also set observabilityMode: true, which stops the GFE waiting on the callout at all. It is not available on a global external Application Load Balancer, where the callout is always inline — there, async shadow is the only thing keeping scoring off the request path.


False positives (legitimate requests flagged)

Symptom: normal API calls appear in the dashboard with high scores

# Check what the inference service returns for a specific payload
kubectl exec -n kysira deployment/kysira-inference -- \
  wget -qO- --post-data='{"request_text":"your payload here"}' \
  --header='Content-Type: application/json' \
  http://localhost:8081/score/all | python3 -m json.tool

The detector field in the response tells you which classifier fired. Common causes:

  • sqli — the SQL injection classifier can be over-eager on SQL keywords in prose. Raise the threshold:

    helm upgrade kysira oci://us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/charts/kysira-platform \
      --namespace kysira --reuse-values \
      --set "kysira-proxy.config.scoreThreshold=0.98"
    
    0.98 significantly reduces false positives at some cost to recall.

  • xss — the regex fires on <script, javascript:, onerror=, and similar. If your app legitimately POSTs HTML content, route those specific paths outside Kysira.

  • nosqli — legitimate Mongo $in/$gte/$lte filter operators in JSON bodies trigger an advisory score (below the kill line by default). They only block in active mode if you've lowered the threshold below 0.7.


Active mode not blocking requests

Symptom: mode is active but malicious requests still reach your app

  1. Confirm the mode change took effect:

    kubectl exec -n kysira deployment/kysira-proxy -- \
      wget -qO- http://localhost:8080/api/mode
    # → {"mode":"active"}
    

  2. Check the dashboard — if action shows shadow_kill instead of active_kill, the pod is still in shadow mode. The mode API sets mode in-memory per pod; use Helm to persist it across restarts: bash helm upgrade kysira oci://us-central1-docker.pkg.dev/cs-poc-uv5os9gxrjsncireus36uzd/kysira-agent-releases/charts/kysira-platform \ --namespace kysira --reuse-values \ --set "kysira-proxy.config.mode=active"

  3. If action shows passed, the score is below the threshold — the request is genuinely not being flagged. Lower the threshold or test with a more obvious payload like ' OR 1=1--.

  4. ext-proc only — the mode API is on the HTTP port (:9090), not the gRPC port, and it is read-only: unlike the proxy, ext-proc fixes its mode at startup from KYSIRA_MODE and has no write path, so a POST returns 405. Changing the mode means redeploying with the new value. Confirm what the running pod is actually set to:

    kubectl exec -n kysira deployment/kysira-ext-proc -- \
      wget -qO- http://localhost:9090/api/mode
    


Dashboard shows no events

Symptom: dashboard loads but the event feed is empty

# Check the proxy is healthy and the dashboard can reach it
kubectl logs -n kysira -l app.kubernetes.io/name=kysira-dashboard

The dashboard proxies /api/ requests to the kysira-proxy service. Upstream connection errors here mean the proxy service name or port is wrong in the dashboard ConfigMap — check the proxyServiceName Helm value matches the proxy service name:

kubectl get svc -n kysira

If the service names look right, check the proxy itself is receiving traffic:

kubectl logs -n kysira -l app.kubernetes.io/name=kysira-proxy | tail -20

HPA not scaling inference

Symptom: inference pod count stays at 1 under load

The HPA requires the Metrics Server. Check if it's installed:

kubectl top pods -n kysira

If this fails, install the Metrics Server:

kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml

Then check the HPA status:

kubectl get hpa -n kysira
kubectl describe hpa -n kysira kysira-inference

The HPA targets 80% CPU by default. If inference is GPU-bound, CPU utilization will be low and the HPA won't trigger — set a fixed replica count or use a custom Prometheus-based metric instead.


Checking metrics directly

# Proxy metrics
kubectl exec -n kysira deployment/kysira-proxy -- \
  wget -qO- http://localhost:8080/metrics | grep kysira_

# ext-proc metrics
kubectl exec -n kysira deployment/kysira-ext-proc -- \
  wget -qO- http://localhost:9090/metrics | grep kysira_extproc_

# Inference metrics
kubectl exec -n kysira deployment/kysira-inference -- \
  wget -qO- http://localhost:8081/metrics | grep kysira_inference_

See Observability for connecting these to Grafana or Datadog.