
Pritam Mazumdar
Senior Devops Engineer
Skills

Bekijk mijn diensten

Portfolio
Werkervaring
Senior Devops Engineer
Qubic • Fulltime
Mar 2023 - Present • 3 yrs 4 mos
Cloud Cost Optimization - Fixed a silent Spot allocation bug that was causing AWS to ignore most instance types, switched to capacity-optimized strategy across 19+ production services - Managed Reserved Instance renewals for RDS, ElastiCache, and OpenSearch — eliminated unplanned on-demand overage costs Infrastructure as Code - Authored and maintained a shared Terraform ECS capacity provider module (v1.0 → v1.3.2) used by all engineering teams - Migrated multiple services from EC2/Beanstalk to ARM Graviton ECS — writing full Terraform from scratch per service - Built ephemeral per-PR environments that spin up on PR open and auto-teardown on merge CI/CD & Reliability - Root-caused and fixed critical race conditions in ECS deployment teardown (ALB drain timing, exit code propagation, retry logic) - Built GitHub Actions workflow suites from scratch for multiple products — Docker → ECR → ECS, hotfix flows, self-hosted runners - Migrated Jenkins from static IAM users to IAM roles, eliminating long-lived credentials Security & Governance - Led org-wide IAM hygiene: removed unnecessary admin access, rotated long-lived keys, hardened DNS and email records - Operationalized GuardDuty for PCI DSS audit readiness with automated alerting and least-privilege roles Observability - Built the ClickStack Observability Platform end-to-end: OpenTelemetry tracing, YAML-based alerting, Python SDK for dashboards-as-code, OTel Collector sidecar on ECS Data Infrastructure - Provisioned ClickHouse cluster on AWS, set up Google SSO, maintained PeerDB replication from Postgres as features shipped
Devops Engineer
Amazon • Fulltime
Sep 2017 - Mar 2023 • 5 yrs 6 mos
Saved $40K/year and 9,457 TB of S3 storage by building an automated audit tool across multiple AWS accounts. Improved API performance by 35% by implementing a cache mechanism with ML-based filtering to handle robotic traffic during peak seasons. Automated service dashboards and monitors using infrastructure-as-code, eliminating manual errors. Enabled fully automated CD pipelines for 12+ services — letting dev teams ship by simply pushing to Git. Reduced operational noise by auditing high-ticket monitors, fixing root causes, and creating SOPs for cross-team workflows. Built a Python-based monitoring system to track 1000+ job flows and surface actionable error reports for on-call teams.
Assistant System Engineer
Tata Consultancy Services • Fulltime
Sep 2014 - Sep 2017 • 3 yrs
TCS — American Express (AMEX) | Global Communication System (GCS) Built monitoring and automation from the ground up for a critical real-time OTP and batch email platform serving AMEX customers. Wrote shell scripts to monitor disk space, CPU, and hung threads across production servers — enabling proactive alerting before outages. Automated job monitoring and built auto-redrive scripts to unstick database records and prevent customer-facing delays. Monitored hundreds of jobs and database record statuses end-to-end. The impact was significant — reduced the team headcount from 12 to 5 in 2 years purely through automation, with zero compromise on reliability.