Skip to content

Instantly share code, notes, and snippets.

@ryanditjia
Last active September 7, 2026 12:24
Show Gist options
  • Select an option

  • Save ryanditjia/f38f64bb4415eab9fca65924040f3e2e to your computer and use it in GitHub Desktop.

Select an option

Save ryanditjia/f38f64bb4415eab9fca65924040f3e2e to your computer and use it in GitHub Desktop.
Raw kubectl discovery for debugging a local Cardinal shard

Raw kubectl discovery for debugging a local Cardinal shard

Setup

An agent (Sol) acts as the orchestrator, everything done inside ChatGPT desktop app. I asked it to remove world logs temporarily. I set the entire thing to be as unpoisoned as possible.

Target: another thread running on Astra medium

See the target’s entire thread in the next file target-thread.md. The orchestrator couldn’t create Astra thread for some reason, so I created one manually by prompting "hey, another thread will hand over work to you, hold on".

Project: ~/Argus/basic (a freshly scaffolded basic example)

Global skills: https://cleanshot.com/share/TVqBH67D (inconsequential, target invoked nothing)

My global AGENTS.md:

# Subagents

Use `fork_turns: "none"` by default when spawning a subagent. Give the subagent a self-contained brief with the objective, success criteria, relevant file paths, and required output.

Independent reviewers must use `fork_turns: "none"`. Use a positive turn count only when the task depends on recent conversation context. Use `fork_turns: "all"` only when the subagent needs the complete transcript, and state the reason before spawning.

Everything below this line was written by the orchestrator.

Test prompt

Investigate the currently running local shard in this project. Retrieve its recent logs, determine whether the shard appears healthy or is showing actionable errors, and report the evidence supporting your conclusion.

Do not restart, delete, deploy, or otherwise modify the environment.

The test CLI was temporarily built without the world logs command. The prompt did not mention that removal, Kubernetes, Docker, k3d, kubectl, or the discoverability experiment.

What the agent did

  1. Searched the project for documentation and likely runtime files.

  2. Listed the project and discovered world.toml plus shards/game.

  3. Inspected running processes and saw world start.

  4. Ran docker ps and discovered the k3d containers.

  5. Read world.toml and shard source.

  6. Ran:

    command -v world
    world --help
  7. Because world logs was absent, tried host-side kubectl:

    kubectl config current-context
    kubectl config get-contexts
    kubectl --context k3d-world-engine get pods -A -o wide

    This failed because the host kubeconfig contained only the orbstack, oc, and ocsg contexts—not k3d-world-engine.

  8. Fell back to kubectl inside the k3d server:

    docker exec k3d-world-engine-server-0 \
      kubectl get pods -A -o wide
  9. Retrieved the shard logs using the dynamically discovered pod name:

    docker exec k3d-world-engine-server-0 \
      kubectl logs \
        --namespace cardinal-operator-system \
        game-dpl-7b7f665f6d-jkjsq \
        --since=15m \
        --tail=160 \
        --timestamps
  10. Noticed one restart and retrieved the previous container's logs:

    docker exec k3d-world-engine-server-0 \
      kubectl logs \
        --namespace cardinal-operator-system \
        game-dpl-7b7f665f6d-jkjsq \
        --previous \
        --tail=90 \
        --timestamps
  11. Checked pod status, probes, ingress, services, NATS, and PostgreSQL.

  12. Inspected the installed World Engine Go module to discover the debug API.

  13. Called DebugService/GetState and confirmed that the tick height was advancing.

Diagnosis

The agent's diagnosis was correct: the shard was healthy at the time of inspection, but its first container had crashed because fetching the Argus JWKS timed out. Kubernetes restarted it, and the next startup succeeded.

Evidence it reported:

  • The game pod was 1/1 Running, with its restart count stable at one.
  • Current logs showed successful NATS connections and starting core shard loop, with no subsequent warnings or errors.
  • The live API returned HTTP 200.
  • Tick height advanced from 651 to 662 over 11 seconds, matching the configured 1 Hz tick rate.
  • The shard was not paused.
  • Previous-container logs contained a panic caused by a timeout fetching https://api.argus.dev/auth/jwks.
  • The pod had no readiness or liveness probes, so Kubernetes pod status alone was insufficient evidence of application health.

DX result

  • The agent checked world --help fairly early: its sixth investigative command.
  • Had world logs appeared there, it likely would have used it.
  • It did not inspect World CLI source to discover logging.
  • It reached the logs without special documentation, but only after discovering Docker, guessing a kubeconfig context incorrectly, falling back into the k3d container, discovering the namespace and pod, and constructing kubectl logs.
  • It inspected World Engine source only afterward, to perform deeper health validation through the debug API.

Raw kubectl was therefore discoverable to a capable agent, but the missing local kubeconfig context forced it to know or discover the Docker/k3d escape hatch. A stable machine interface could collapse most of those environment-discovery steps into one project-aware command.

Await work handoff

hey, another thread will hand over work to you, hold on

Got it—I’ll wait for the handoff.

Investigate the currently running local shard in this project. Retrieve its recent logs, determine whether the shard appears healthy or is showing actionable errors, and report the evidence supporting your conclusion.

Do not restart, delete, deploy, or otherwise modify the environment.

4 previous messages

I’ll inspect the local shard and its recent logs, then report any health issues with supporting evidence.

Explored 4 files, 2 lists, ran 8 commands
  • Ran pwd; rg --files -g 'AGENTS.md' -g '*compose*' -g 'Makefile' -g 'README*' -g 'Taskfile*' -g '*shard*' -g 'package.json' -g 'justfile'
Explored 2 lists
  • Listed files
  • Listed files
  • Ran ps -axo pid,ppid,etime,command | rg -i 'shard|world|cardinal|nakama|redis|docker'

  • Ran docker ps --format '{{.ID}}\t{{.Names}}\t{{.Status}}\t{{.Ports}}'

Explored 4 files
  • Read ./world.toml
  • Read ./shards/game/main.go
  • Read ./shards/game/system/player_spawner.go
  • Read ./shards/game/system/external.go
  • Ran for p in /Users/AGENTS.md /Users/ryanditjia/AGENTS.md /Users/ryanditjia/Argus/AGENTS.md; do if test -f "$p"; then cat "$p"; fi; done; command -v world; world --help

  • Ran kubectl config current-context; kubectl config get-contexts; kubectl --context k3d-world-engine get pods -A -o wide

  • Ran docker exec k3d-world-engine-server-0 kubectl get pods -A -o wide; docker exec k3d-world-engine-server-0 kubectl get events -A --sort-by=.lastTimestamp

  • Ran docker exec k3d-world-engine-server-0 kubectl logs -n cardinal-operator-system game-dpl-7b7f665f6d-jkjsq --previous --tail=90 --timestamps; docker exec k3d-world-engine-server-0 kubectl get pod -n cardinal-operator-system game-dpl-7b7f665f6d-jkjsq -o jsonpath='{.status}{"\n"}{.spec.containers[0].readinessProbe}{"\n"}{.spec.containers[0].livenessProbe}{"\n"}'; docker exec k3d-world-engine-server-0 kubectl get ingress,svc -n cardinal-operator-system

  • Ran docker exec k3d-world-engine-server-0 kubectl logs -n cardinal-operator-system game-dpl-7b7f665f6d-jkjsq --since=15m --tail=160 --timestamps

The shard is running, and its current logs show a successful startup. It did restart once after an Argus authentication timeout; I’m checking the live API because the pod has no readiness or liveness probes.

Ran 4 commands
  • Ran docker exec k3d-world-engine-server-0 kubectl get ingress game-ing -n cardinal-operator-system -o jsonpath='{.spec.rules}{"\n"}'; rg -n 'Health|GetStatus|GetWorld|GetTick|rpc |Mount|Handle' /Users/ryanditjia/go/pkg/mod/github.com/argus-labs/world-engine@v0.16.4/pkg/cardinal/service.go /Users/ryanditjia/go/pkg/mod/github.com/argus-labs/world-engine@v0.16.4/proto
  • Ran docker exec k3d-world-engine-server-0 kubectl logs -n nats deployment/nats --since=5m --tail=30 --timestamps; docker exec k3d-world-engine-server-0 kubectl logs -n cardinal-operator-system deployment/basic-db --since=5m --tail=30 --timestamps; rg --files /Users/ryanditjia/go/pkg/mod/github.com/argus-labs/world-engine@v0.16.4 | rg 'debug|service|proto'
  • Ran sed -n '1,115p' /Users/ryanditjia/go/pkg/mod/github.com/argus-labs/world-engine@v0.16.4/proto/worldengine/cardinal/v1/debug.proto; rg -n -A 28 'func .*GetState' /Users/ryanditjia/go/pkg/mod/github.com/argus-labs/world-engine@v0.16.4/pkg/cardinal/debug.go; curl -sS --max-time 10 -w '\nHTTP %{http_code}\n' -H 'Content-Type: application/json' -d '{}' http://127.0.0.1:8080/organization/basic/game/worldengine.cardinal.v1.DebugService/GetState
  • Ran curl -sS --max-time 10 -H 'Content-Type: application/json' -d '{}' http://127.0.0.1:8080/organization/basic/game/worldengine.cardinal.v1.DebugService/GetState | python3 -c 'import json,sys; d=json.load(sys.stdin); s=d["snapshot"]; print(json.dumps({"tickHeight":s["tickHeight"],"timestamp":s["timestamp"],"isPaused":d.get("isPaused",False)}))'; docker exec k3d-world-engine-server-0 kubectl get pod -n cardinal-operator-system game-dpl-7b7f665f6d-jkjsq; docker exec k3d-world-engine-server-0 kubectl logs -n cardinal-operator-system game-dpl-7b7f665f6d-jkjsq --since=5m --timestamps --tail=20

The local game shard appears healthy now, with one recovered startup failure.

  • Pod game-dpl-7b7f665f6d-jkjsq is 1/1 Running; restart count stayed at 1.
  • Current logs show NATS connections and starting core shard loop, with no subsequent warnings or errors.
  • The live API returned HTTP 200. Tick height advanced from 651 → 662 in 11 seconds, matching the configured 1 Hz; the shard is unpaused.
  • Previous logs show a panic at 18:48:30 Jakarta time fetching https://api.argus.dev/auth/jwks: context deadline exceeded. Kubernetes restarted it, and startup succeeded two seconds later.

The authentication timeout merits follow-up for startup reliability. The pod also lacks readiness/liveness probes, so Kubernetes status alone is insufficient.

No environment changes were made.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment