Senior Staff Engineer Bengaluru, India — remote-first

SumanPatra

Platform Engineering · Observability · IDP · AI-native tooling

I build the platforms other engineers build on — Internal Developer Platforms, cloud-native observability, and now AI-native developer tooling. 10+ years across enterprise banking, e-commerce, and CDN-scale engineering. Remote-first.

10+ yrs
Platform & infra engineering
100 TB/day
Managed-Loki logs · 50+ clusters · 40 teams
$1.4M/yr
Saved by replacing observability SaaS with an owned stack
−30% time
Delivery time · AI SDLC harness across ~500 developers
KubernetesAWSGo TerraformOpenTelemetryGrafana Loki ClickHouseIDPArgo CD CrossplaneObservabilityAI agents

About

Senior Staff Platform Engineer with 10+ years across platform, infra, and observability. Founding engineer on two production IDPs. I run the whole arc — evaluate and select core technology, form the cross-functional team, ship to production, and drive org-wide adoption. Work spans enterprise banking, social commerce, and CDN-scale developer tooling — same craft: paved roads, self-service, and owned platforms at scale.

Experience

Nagarro — Senior Staff Platform Engineer, Platform & Observability

Mar 2023 – Present
  • Migrated ~1.5M pages org-wide to Atlassian Cloud in two quarters — and rebuilt the corpus for humans and agents — ~35–40K engineering documents re-authored into a hierarchical wiki, with CI/CD gates mapping every commit back to its documentation so knowledge can't drift from code.
  • Built the SDLC agent harness on top of it — ~500 developers, ~30% faster delivery — a service catalog that assembles multi-microservice context on demand, a RAG layer over the curated codebase, and purpose-built skills for code, review, quality, Jira ticketing and ADRs. Carries work from brainstorm → Jira → architecture → build → test plan → release → debug.
  • Onboarded 38 teams onto a from-scratch IDP, cutting new-service setup from ~3 days to 15 min (first commit & deploy) — golden-path scaffolding (Java/Node/Go/Python), Tekton + Argo CD + OpenShift, self-service data abstractions on Crossplane.
  • Scaled managed Loki to 100 TB/day across 50+ clusters and 40 teams, cut shared-infra cost ~60% via a single multi-tenant model serving a 100-person SRE org.
  • Replaced third-party observability SaaS with a fully owned stack, saving $1.4M/year — chose ClickHouse + ClickStack + HyperDX over SigNoz/OpenSearch for compression & maintainability.

Nutanix — Member of Technical Staff 4

Feb 2022 – Apr 2023
  • Lifted on-call efficiency +20% and cut response time −30% — led the end-to-end PagerDuty replacement (eval → POC → 6-engineer team → prod in under 3 quarters).
  • Delivered 100K+ monthly alerts at sub-5s latency to 200+ engineers — rotation & notification microservices on AWS (Spring Boot, RabbitMQ, Slack, Twilio).

Akamai Technologies — Senior Software Engineer

Jun 2020 – Feb 2022
  • Cut Git-admin support tickets 40% — NLP self-service chatbot for a company-wide Perforce→Git migration.
  • Governed repo access for 2,500+ developers across Bitbucket, Perforce & GitHub, plus an engineering-metrics dashboard adopted across 7 release trains (~30–35 teams).

Flabber Technologies (alippo.com) — Founding Member

Jun 2018 – May 2020
  • Enabled 100K+ SKUs across 3 marketplaces (Amazon, Flipkart, Meesho) via one integration — catalog fan-out, order routing, inventory sync (Java, Spring Boot, Kafka, Redis, AWS).
  • Scaled the platform to 1,000 orders/day and 1.5M weekly active users — cataloging (Elasticsearch), Redis hot-path, CDN media, order management, and shipment modules.

Akamai Technologies — Software Engineer II

May 2015 – May 2018
  • Decomposed JSF monoliths into AngularJS SPAs + independently deployable Java services; set REST/API conventions adopted across teams.

Selected Work

Internal Developer Platform (IDP)

From-scratch paved road: scaffold → build → deploy → run, self-service across four languages.

Developer Portal self-service Golden-path Java·Node·Go·Python CI/CD Tekton · Argo CD OpenShift Kubernetes Self-service data: Redis · MySQL · Elasticsearch

Managed Observability Platform

Multi-tenant logs at 100 TB/day + a fully owned logs/metrics/traces stack replacing SaaS.

50+ K8s clusters 40 teams · multi-tenant OpenTelemetry collect · route Managed Loki 100 TB/day logs ClickHouse+HyperDX metrics · traces SRE org 100 people

Escalation-Management Platform (PagerDuty replacement)

Owned end-to-end; +20% on-call efficiency, −30% response time, 100K+ alerts/mo at sub-5s.

Incident source alerts in Rotation+Escalation Spring Boot·RabbitMQ Notification svc AWS · sub-5s Slack Twilio SMS Voice call

Open source

Chakra — AI-native deployment platform

~27K lines of Go

An open-source platform that abstracts operations, not just servers. Push code to a branch; the Brain infers every infrastructure dependency, generates the Dockerfile and Kubernetes manifests, provisions databases and caches, and delivers a running app at a live URL — no YAML written by hand. Built on Pulumi-provisioned clusters with Argo CD, Traefik, and cert-manager, with logs, metrics, and traces native to the platform from the first deploy.

GoKubernetesArgo CDGitOps PulumiTraefikcert-managerHelm

Vajra

~16K lines of Rust

Provable governance for AI coding agents — a Rust binary that enforces rules through hooks and fails closed, rather than suggesting them in a prompt. Runs a governed multi-agent SDLC pipeline across eight stations (Analyst → Architect → Planner → Coder → QA → Demo-er → Releaser → Reviewer) with delta-tracked handoffs. Every verdict is attested as sha256(prompt‖diff) and chained into a tamper-evident append-only ledger, and each session emits an honest cost receipt.

Dogfooded, not demoed: the workflow has driven 260+ governed sessions across 8 projects since March 2026 — including the platform work above.

RustPolicy enforcementAttestationAppend-only ledger Developer toolingAI agentsCost metering

devx

Crossplane

Self-service data infrastructure as Crossplane compositions — PostgreSQL and Redis claimed by developers as first-class resources, with provider-helm and provider-sql wired underneath, plus a small CLI to drive it. The paved-road pattern behind self-service databases on an IDP.

CrossplaneKubernetesHelmPostgreSQL RedisPlatform engineering

Kreeda & inini.in

Live in production

Kreeda is a multi-tenant operations platform where agents run the day-to-day and the founder acts as the board. inini.in is tenant #1 — an agent-operated desk publishing one approved piece a day, live today. Multi-tenancy, approval gates, and scheduled autonomous work, running against real traffic rather than a demo.

Multi-tenantAgent orchestrationSchedulingProduction

Writing

An engineer's guide to thinking like the machine — intuition-first, no calculus wall.

Sutra In development

An AI SaaS factory for spinning up production micro-SaaS, fast.

Essays Ongoing

Notes on AI engineering, developer productivity, and platform strategy.