T.Anandakkoomar
Folio I · Platform & Site Reliability Engineering

Works in production.

I build the Kubernetes platforms, pipelines and observability that let engineers deploy without fear, and I run this page the same way.

Folio II

About

Most outages are not exotic. They come from slow, frightening deploys, from alerts nobody trusts, and from systems only one person understands. My work is removing those conditions.

I like boring reliability: deploys that are routine, dashboards that answer real questions, and runbooks written for whoever is paged at three in the morning. Everything I build is defined in code and reviewed like code, this page included.

Instruments of the trade: Kubernetes, Terraform, GitOps, GitHub Actions, Prometheus, Grafana, SLOs, incident response.

Folio III

Results

6min
deploy time, formerly 45 min

A pipeline people use

Problem
30 services, a manual checklist, weekly batches.
Method
Shared CI template, cached builds, canary with auto-rollback.
Result
Several deploys a day instead of one a week.
−60%
on-call pages, ≈23 to ≈9 a week

Alerts on symptoms

Problem
Hundreds of threshold alerts, mostly ignored.
Method
SLOs per user journey; page only on error-budget burn.
Result
Fewer pages, no missed incidents.
−38%
compute spend

Right-sized clusters

Problem
Sized for peak, all year round.
Method
Measured requests vs limits, autoscaling, spot for batch work.
Result
Same latency, much smaller bill.
0
hand-built resources, formerly ≈400

Everything as code

Problem
Cloud resources made by hand, understood by few.
Method
Imported into Terraform module by module; plans reviewed on every PR.
Result
Nightly drift detection, nothing hand-made left.
Folio IV

Status

99.98%

availability over the last 30 days, against an objective of 99.9 %

Measurements
WindowUptimep95 latency
24 hours100.00 %142 ms
7 days100.00 %151 ms
30 days99.98 %158 ms

All systems operational. Read every five minutes by an independent observer outside the machine; full record.

Folio V

Stack

Sketch of the server cabinet. Scrolling separates it into labelled parts: the lid (Cloudflare), the gear (Terraform), three drawers (cloudflared, nginx, the update timer) and the power unit (the Hetzner server). lid: Cloudflare gear: Terraform cloudflared nginx update timer power: Hetzner CAX11

Scroll to take the machine apart; scroll back to rebuild it.

Parts of the machine
PartOffice€/mo
CloudflareDNS, TLS, WAF, cache, tunnel0.00
Hetzner CAX112 vCPU ARM server, no inbound ports4.50
cloudflaredopens the tunnel from inside0.00
nginxserves this page, read-only0.00
systemd timerfetches each new image within 60 s0.00
GitHub Actionsbuilds, checks, publishes, verifies0.00
Terraformdescribes every part above0.00
Upptimethe independent observer0.00
Total€4.50

Next folio: the cabinet becomes three, joined into a Kubernetes cluster with GitOps, metrics and published objectives.