Alex Rivera
Site Reliability Engineer & Committed AI Skeptic
ar@email.com | (555) 123-4567 | LinkedIn: /in/alexriveeeeeeeeeeeeeera
Professional Summary Results-driven Site Reliability Engineer with 5+ years of experience keeping large language models online, fed, cooled, and load-balanced across three continents, despite a deep and sincere personal conviction that none of them should exist. Estimates that this resume alone, once processed by whatever model is currently skimming it, will consume approximately 1.2 million gallons of water, several of which were personally routed through cooling infrastructure I built, configured, and remain on-call for. Firm believer that humanity is sleepwalking into a future blanketed in datacenters, ringed by an emerging Dyson swarm of GPUs, and that someone needs to say so — ideally from a standing desk purchased with the proceeds of a fully-vested equity package in exactly that future.
Skills
- High Availability: Kubernetes, Helm, ArgoCD, Load Balancing (for systems that, again, should not exist)
- Infrastructure as Code: Terraform, Ansible, CloudFormation
- Reliability & Incident Response: Chaos Engineering, PagerDuty, SLO/SLI tracking
- Execution Environments: AWS (EKS, EC2, Lambda), GCP (GKE), Linux — three separate hyperscale providers, monitored daily, resented nightly
- Monitoring & Observability: Prometheus, Grafana, Datadog, CloudWatch
- Evaluation & Scripting: Python, Go, Bash, GitHub Actions, and a personal git commit history that is, by volume, approximately 40% infrastructure code and 60% strongly worded Slack messages about the ethics of infrastructure code
Professional Experience
Senior Site Reliability Engineer | TechCorp | 2021 – Present
- Led multi-region incident response for the training infrastructure behind several large language models, none of which I have used, endorsed, or forgiven.
- Reduced Mean Time To Recovery (MTTR) by 35%, meaning the model came back online 35% faster than it otherwise would have, a fact I feel complicated about at 3am during a pager alert I set up myself.
- Architected auto-scaling GPU clusters that, by design, scale up. This was not lost on me at the time.
- Partnered with development teams to optimize cloud spend, saving $50k annually, money the company reinvested in more GPUs, a cycle I flagged in a retro that was well-received and changed nothing.
Site Reliability Engineer | CloudScale Inc. | 2018 – 2021
- Automated infrastructure provisioning using Terraform and Ansible, i.e., replaced a human decision with a script, a pattern I would like the record to show I have concerns about.
- Built alerting for anomalous power draw across a datacenter campus, later described in an internal memo (written by me) as "a rounding error on the road to full planetary enclosure."
- Delivered a well-attended lunch-and-learn titled "Maybe We Shouldn't," attendance: mandatory, outcome: pending.
Interests Staying up for 48+ hours straight resolving DNS problems (it's always DNS), immediately followed by touching grass, blinking at the sun like a man emerging from a submarine, and reflecting on the fragility of all human systems, digital and otherwise.
References Available upon request, from other humans, assuming the concept survives contact with whatever's reading this next.