Console style
// staff_site_reliability_engineer

Kiran Mehta

SRE who treats an outage as a design review that arrived late
Get in touch
Based in Toronto, ON
Based in Toronto, ON
99.99%
Uptime held
4 min
Median detect
70%
Fewer pages
how_i_work.list

Alert on symptoms

Pages fire on what the customer feels, not on what a dashboard happens to measure.

Blameless, but specific

The postmortem names the system that failed and the change that stops it, never the person on call.

Automate the second time

Does it by hand once, writes the runbook, then deletes the runbook by automating it.

pages_per_week,_by_quarter.chart
Jan
52
Mar
63
May
71
Jul
82
Sep
91
Dec
100
career_path.log
[2023 — present] Staff SRE @ Kelsey Pay
[2020 — 2023] Senior SRE @ Latchbox
[2018 — 2020] Backend engineer @ Vero Logistics
stack_and_tooling.json
Kubernetes
95%
Terraform
91%
Go
87%
Observability
83%
Incident command
79%
Capacity planning
75%
systems_i_have_owned.json
Payments tier
95%
Dispatch service
91%
Global CDN edge
87%
On-call programme
83%
what_people_say.txt

Kiran rewrote our alerting and the on-call rotation stopped burning people out. Retention on that team went up.

Sam Ibarra, Director of Engineering, Kelsey Pay

His postmortems are the only ones I read end to end. They say what broke and what changes, and nothing else.

Wen Zhao, Principal Engineer, Latchbox

Get in touch

Architecture writing samples and references available.

# Kiran Mehtaconsole.log('built with pitchpage')
Like this page? Build your own at pitchpage.co