Skip to content

Instantly share code, notes, and snippets.

View jmealo's full-sized avatar

Jeff Mealo jmealo

  • Gisual
  • Pennsylvania
  • 07:00 (UTC -04:00)
View GitHub Profile
@jmealo
jmealo / README.md
Last active September 1, 2026 03:39
Run legacy 32-bit Adobe AIR macOS games (e.g. Detective Grimoire, 2014) on Apple Silicon / modern macOS. Bring your own legally purchased copy.

Reviving 32-bit Adobe AIR macOS games on Apple Silicon

A script that rebuilds a legacy Adobe AIR game you already own so it runs on modern macOS (Apple Silicon and Intel).

This ships no game content and no Adobe/Wipro code. It only transforms files already on your machine. You supply your own legally purchased copy.

Tested with Detective Grimoire (SFB Games, 2014). It may work with other Adobe AIR apps, but that is untested. Nothing is hardcoded, since the script

@jmealo
jmealo / cnpg-switchover-runbook.md
Last active April 4, 2026 03:09
CNPG pg-common Primary Switchover Runbook — Staging, Demo, Production

CNPG pg-common Primary Switchover Runbook

Date: 2026-04-04 Objective: Migrate pg-common primary to newly provisioned nodes across staging, demo, and production, then decommission old-node instances to reach a 3-replica cluster per environment.

Current State

Staging (azure-staging / staging-pg-common)

| Instance | Role | Node | Node Age |

@jmealo
jmealo / otel-metrics-rollout.md
Created March 22, 2026 19:11
OTel Metrics Rollout — Monday Checklist for Robin

OTel Metrics Rollout — Monday Checklist

Background

The prometheus_client Python library uses threading.Lock per label combination on every .labels().observe() call. This caused OOM crashes (700+ MiB → OOMKill) on aaa-api in staging. The fix replaces the prometheus_client backend with OpenTelemetry via a PROMETHEUS_BACKEND=otel env var toggle in gisual-prometheus-clients.

Validated in staging since 2026-03-22: aaa-api running at 162 MiB (under 368 MiB limit), zero restarts, zero 500 errors, all metrics present in /metrics output.


@jmealo
jmealo / 2026-03-16-rca-missing-intel-requests-created.md
Last active March 16, 2026 23:16
RCA: Missing intel_requests.created — Corrected Analysis (ENG-3349)

RCA: Missing intel_requests.created (ENG-3349) — Corrected

Status PROVISIONAL — ASU-specific loss confirmed, root cause unknown
Supersedes Previous RCA
Incident date 2026-03-10
Analysis date 2026-03-16
Reverted intel-requests-api !48 (batch_id change from !47)
@jmealo
jmealo / 2026-03-14-rca-intel-api-amqp-cascade.md
Last active March 14, 2026 23:08
RCA: Intel-API AMQP Cascade Failure — CPU Limit Regression (2026-03-14)

RCA: Intel-API AMQP Cascade Failure — CPU Limit Regression

Date: 2026-03-14 Environment: Production Severity: Critical (P1) Duration: ~23 hours (2026-03-14 00:15 UTC — ongoing) Status: IN PROGRESS — fix identified, rolling restart underway

Summary

@jmealo
jmealo / staging-deploy-status.md
Created March 14, 2026 15:16
Staging Deploy Status — 2026-03-14 — Metrics Changes

Staging Deploy Status — 2026-03-14

Summary

Deployed 25 backend services to staging to test metrics changes. Found and fixed 5 bugs introduced by library version mismatches. All services are now running healthy.

Libraries Published

Library Version Fix
@jmealo
jmealo / search-retry-feeder-crashloop-rca.md
Last active March 11, 2026 16:37
RCA: search-retry-feeder CrashLoopBackOff (Production, 2026-03-11)

RCA: search-retry-feeder CrashLoopBackOff (Production)

Date: 2026-03-11 Reported by: Diego via Slack/Datadog Investigated by: Jeff Mealo (with Claude Code) Severity: Low (self-recovering, no data loss) Status: Root cause identified, fix committed and pending deploy

Summary

@jmealo
jmealo / notification-sender-42-red-herrings.md
Created March 6, 2026 16:47
42 red herrings across 2 days of notification-sender queue backup investigation (March 5-6, 2026)

42 Red Herrings: Notification-Sender Queue Backup (March 5-6, 2026)

Two incidents, one root cause chain, 42 documented false signals across 2 days of investigation.

Actual root cause: AAA API at 8 replicas (some crash-looping) couldn't serve permissions queries from intel-requests-api fast enough. This caused cascading request queuing through intel-requests-api and incidents-api, starving notification-sender of API capacity. SQL queries were sub-millisecond throughout.


March 5: 29 Red Herrings (5-hour investigation)

@jmealo
jmealo / notification-sender-explain-queries-2026-03-06.md
Last active March 6, 2026 08:53
EXPLAIN ANALYZE queries for notification-sender bottleneck — real prod IDs from 2026-03-06

EXPLAIN ANALYZE Queries for Notification-Sender Hot Path

Source: Production logs, 2026-03-06 08:44 UTC Charter org_id: 9de5d801-235a-4451-8c89-d2c3974c71e8

All queries use real IDs extracted from prod notification-sender logs.

WARNING: Run these inside a BEGIN; ... ROLLBACK; transaction on a read replica if possible. EXPLAIN ANALYZE actually executes the query.

@jmealo
jmealo / notification-sender-bottleneck-analysis-2026-03-06.md
Last active March 6, 2026 15:31
notification-sender bottleneck deep dive — recursive CTE, API call chain, optimization targets (2026-03-06)