Skip to content

Instantly share code, notes, and snippets.

@Excedrin
Created August 23, 2026 20:28
Show Gist options
  • Select an option

  • Save Excedrin/6cc41af0ce79a08cb0d0cb0c3496aeb4 to your computer and use it in GitHub Desktop.

Select an option

Save Excedrin/6cc41af0ce79a08cb0d0cb0c3496aeb4 to your computer and use it in GitHub Desktop.
resume.md

Alex Rivera

Site Reliability Engineer & Committed AI Skeptic

ar@email.com | (555) 123-4567 | LinkedIn: /in/alexriveeeeeeeeeeeeeera

Professional Summary Results-driven Site Reliability Engineer with 5+ years of experience keeping large language models online, fed, cooled, and load-balanced across three continents, despite a deep and sincere personal conviction that none of them should exist. Estimates that this resume alone, once processed by whatever model is currently skimming it, will consume approximately 1.2 million gallons of water, several of which were personally routed through cooling infrastructure I built, configured, and remain on-call for. Firm believer that humanity is sleepwalking into a future blanketed in datacenters, ringed by an emerging Dyson swarm of GPUs, and that someone needs to say so — ideally from a standing desk purchased with the proceeds of a fully-vested equity package in exactly that future.

Skills

  • High Availability: Kubernetes, Helm, ArgoCD, Load Balancing (for systems that, again, should not exist)
  • Infrastructure as Code: Terraform, Ansible, CloudFormation
  • Reliability & Incident Response: Chaos Engineering, PagerDuty, SLO/SLI tracking
  • Execution Environments: AWS (EKS, EC2, Lambda), GCP (GKE), Linux — three separate hyperscale providers, monitored daily, resented nightly
  • Monitoring & Observability: Prometheus, Grafana, Datadog, CloudWatch
  • Evaluation & Scripting: Python, Go, Bash, GitHub Actions, and a personal git commit history that is, by volume, approximately 40% infrastructure code and 60% strongly worded Slack messages about the ethics of infrastructure code

Professional Experience

Senior Site Reliability Engineer | TechCorp | 2021 – Present

  • Led multi-region incident response for the training infrastructure behind several large language models, none of which I have used, endorsed, or forgiven.
  • Reduced Mean Time To Recovery (MTTR) by 35%, meaning the model came back online 35% faster than it otherwise would have, a fact I feel complicated about at 3am during a pager alert I set up myself.
  • Architected auto-scaling GPU clusters that, by design, scale up. This was not lost on me at the time.
  • Partnered with development teams to optimize cloud spend, saving $50k annually, money the company reinvested in more GPUs, a cycle I flagged in a retro that was well-received and changed nothing.

Site Reliability Engineer | CloudScale Inc. | 2018 – 2021

  • Automated infrastructure provisioning using Terraform and Ansible, i.e., replaced a human decision with a script, a pattern I would like the record to show I have concerns about.
  • Built alerting for anomalous power draw across a datacenter campus, later described in an internal memo (written by me) as "a rounding error on the road to full planetary enclosure."
  • Delivered a well-attended lunch-and-learn titled "Maybe We Shouldn't," attendance: mandatory, outcome: pending.

Interests Staying up for 48+ hours straight resolving DNS problems (it's always DNS), immediately followed by touching grass, blinking at the sun like a man emerging from a submarine, and reflecting on the fragility of all human systems, digital and otherwise.

References Available upon request, from other humans, assuming the concept survives contact with whatever's reading this next.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment