About me

I'm a Senior Platform Engineer based in Dhaka, Bangladesh, specializing in cloud-native infrastructure, DevOps automation, and internal developer platforms. I currently work at Brain Station 23, where I lead platform engineering across fintech and banking products.

I love building simple, scalable, and easy-to-operate systems. My work spans designing production-grade Kubernetes clusters on AWS EKS, building GitOps pipelines with Flux v2, embedding DevSecOps practices, and defining SLIs/SLOs to keep reliability commitments. I enjoy mentoring engineers, solving complex infrastructure challenges, and enabling product teams to ship faster with less toil.

What I do

  • platform icon

    Platform Engineering

    Designing and operating internal developer platforms on Kubernetes that let product teams self-serve deployments and eliminate infrastructure bottlenecks.

  • cloud icon

    Cloud Infrastructure

    Architecting production-grade AWS environments using CDK and Terraform — VPC, EKS, IAM, S3, CloudWatch — across multiple availability zones with HA and DR built in.

  • devsecops icon

    DevSecOps

    Shifting security left with Trivy, Vault, Keycloak, and Boundary — while also building backend services and integrations in Java and Spring Boot.

  • observability icon

    Observability & SRE

    Building symptom-based alerting with OpenTelemetry, Grafana Loki, Tempo, and Prometheus. Defining SLIs and SLOs, managing error budgets, and running on-call incident response.

> ask me anything...

Experience

Career Path
LVL 2

Senior Software Engineer I

[DEVOPS] ● ACTIVE
Brain Station 23 Aug 2024 — Present
🔗 500 services ✅ 99% uptime 🛡 100% RBAC 📡 95% pipelines
  • Integrated mobile app CI/CD covering 95% of service pipelines across 15 projects (~500 services), and led DevSecOps adoption by embedding vulnerability scanning (Trivy, SonarQube) into pipelines.
  • Owned on-call incident response for platform and infrastructure, leading L3 escalation and resolution across production environments while contributing to 99% uptime.
  • Defined and tracked SLIs/SLOs for critical platform services, using error budgets to prioritize reliability work and balance feature delivery.
  • Drove toil reduction by identifying and automating repetitive operational tasks across teams, freeing engineering capacity for higher-value improvements.
  • Strengthened security posture by implementing RBAC policies for 100% of users and conducting DevSecOps vulnerability analysis across services.
  • Drove pre-sales activities delivering infrastructure cost estimations, cloud architecture proposals, and technical consulting for client engagements.
  • Served as technical leader and mentor on cloud-native best practices, establishing engineering standards and accelerating team onboarding.
LVL 1

Associate Software Engineer

[DEVOPS]
Brain Station 23 Nov 2021 — Jul 2024
⚡ 80% faster deploys 📉 20% latency cut
  • Streamlined DevOps processes, achieving an 80% reduction in development and operations lifecycle time for Fintech DEV/UAT cluster.
  • Conducted load testing and performance analysis, identifying bottlenecks that reduced latency by 20% after optimization.
  • Designed and developed event-driven microservices with Spring Boot, streamlining data flow, scalability, and fault tolerance.
  • Researched and developed Hyperledger Fabric-based solutions, improving efficiency in supply chain ecosystems.
Education
EDU

Islamic University of Technology

[SOFTWARE ENG]
Dhaka, Bangladesh Jan 2020 — Jul 2024
🎓 B.Sc. Software Engineering ⭐ CGPA 3.57 / 4.00

Skills

Cloud Platforms

AWSGCPAzureHuawei Cloud
Proficiency
70%

Container Orchestration

KubernetesHelmKustomizeFlux v2 KarpenterDocker
Proficiency
72%

CI/CD & GitOps

JenkinsArgo WorkflowsFastlaneBuildpack GiteaJFrog ArtifactoryGitHub Actions
Proficiency
68%

Infrastructure as Code

AWS CDKTerraformCloudFormationAnsible
Proficiency
65%

Observability & SRE

OpenTelemetryPrometheusGrafana LokiGrafana Tempo CloudWatchPMM
Proficiency
67%

Security & Access

HashiCorp VaultKeycloakBoundary TrivyHarborSonarQubeModSecurity
Proficiency
62%

Networking & Ingress

CiliumCalicoNGINXAPISIX ModSecurity WAFALBNAT Gateway
Proficiency
58%

Data & Messaging

PostgreSQLPatroniStrimzi KafkaDebezium RedisMinIOMongoDB
Proficiency
55%

Programming Languages

JavaPythonJavaScriptTypeScript CSQLBash
Proficiency
60%

Application Frameworks

Spring BootReactNode.js
Proficiency
50%

Portfolio

  • Fintech Internal Developer Platform

    Platform & Operations Engineer
    Oct 2023 — Present
    K8STerraformArgoFastlane BuildpackSonarQubeTrivyHarbor BoundaryKeycloakVaultJFrogNGINX
    • Designed and built an internal developer platform on Kubernetes with Cilium networking, enabling 15+ product teams to self-serve deployments and eliminating infrastructure bottlenecks across DEV/UAT environments.
    • Integrated Jenkins, Argo Workflows, Fastlane, Buildpack, Helm, and GitHub Actions to automate deployments across ~500 services, achieving an 80% reduction in delivery cycle time and eliminating toil from manual release processes.
    • Established platform-wide golden paths: 100% PR-based code quality gates (Gitea + SonarQube), automated container vulnerability scanning (Trivy + Harbor), and symptom-based alerting with OpenTelemetry, Grafana Loki, Tempo, and Prometheus.
    • Secured all services with zero-standing-credentials using HashiCorp Vault for dynamic secrets, Keycloak for SSO, and Boundary for RBAC-enforced infrastructure access, enforcing least-privilege across the platform.
  • ZooberPay

    DevOps Engineer
    Mar 2026 — Present
    AWSFluxKarpenterPatroni Strimzi KafkaRedisOpenTelemetry APISIXSpring BootJava
    • Architected a production-grade fintech platform on AWS EKS across 3 AZs using AWS CDK (TypeScript) — 8 independently deployable stacks, 15+ reusable constructs — operated via Flux v2 GitOps managing 17 app workloads with per-component Kustomizations and tight prune scopes.
    • Deployed self-managed HA data tier using Patroni PostgreSQL, Strimzi Kafka, and Redis Failover across 3 AZs with gp3 StorageClasses, achieving sub-minute RPO via pgBackRest WAL archiving to S3 with cross-region replication; validated through controlled failure injection and restore drills.
    • Defined SLIs and SLOs for payment platform services, using error budgets to govern release velocity and prioritize resilience improvements over feature work during budget exhaustion.
    • Developed and integrated a bKash payout system for ZooberPay, enabling seamless mobile financial service disbursements for platform users.
    • Enforced infrastructure compliance using AWS CDK automated provisioning guardrails and CloudTrail, maintaining 100% audit trail coverage across all 8 stacks and AWS resources.
    • Built symptom-based alerting via OpenTelemetry Collector shipping to CloudWatch Logs (6-month retention) with 5 EKS control-plane log streams, reducing alert noise and improving on-call signal fidelity.
  • Revo Global

    DevOps Engineer
    Nov 2025 — Present
    PythonPlaywrightJavaSpring BootSQS
    • Architected a cloud-native browser automation platform using Python Playwright, containerized with Docker (2GB shared memory, VNC/noVNC remote viewing), and deployed as Kubernetes Jobs with persistent browser profiles to scrape deep web data sources at scale.
    • Built a Java Spring Boot backend exposing scraped data through REST APIs, dynamically orchestrating Kubernetes Jobs via Mustache-templated manifests and processing events across 6 AWS SQS FIFO queues with exponential backoff retry, Dead Letter Queue alerting, and 3-attempt fault tolerance.
    • Implemented a manual interruption and monitoring system integrated through Linux window managers and VNC, enabling real-time operator oversight and controlled intervention in automated browser sessions.
    • Secured the platform with HashiCorp Vault and integrated Keycloak v26 for multi-tenant SSO with a multi-layer RBAC hierarchy, caching authorization decisions in Redis for sub-millisecond access control.
    • Provisioned full AWS infrastructure using CDK v2 (TypeScript) across UAT and Production environments, deploying a Kubernetes cluster with Cilium CNI and OpenTelemetry observability pipeline.
  • Oct 2024 — Present
    Huawei CloudBashRundeckKafka MinIORedisJenkinsOpenTelemetry Grafana StackPlaywrightSpock
    • Engineered and deployed 2 environments (UAT and Production) including servers, networks, load balancers, and WAF on Robi Huawei Cloud, supporting a PLC-grade internet banking platform.
    • Eliminated manual server maintenance by implementing full automation via Bash scripts and Rundeck, reducing synchronization task execution time from hours to minutes.
    • Deployed Kafka Raft, MinIO, Redis, and application services in 3-node distributed mode with end-to-end TLS on bare-metal servers, achieving zero single points of failure.
    • Developed solutions to integrate reconciliation systems with MFS platforms, streamlining financial data synchronization across banking and mobile financial services.
    • Developed integration tests for regression using Playwright and Spock, ensuring end-to-end reliability across critical banking workflows.
    • Optimized the event messaging architecture, improving system throughput and reducing processing latency by 50%.
  • City Remittance

    DevOps Engineer
    Jun 2023 — Present
    Spring BootK8SCalicoKafka RedisMinIOPrometheusLokiNGINX
    • Deployed and maintained the application across 2 clusters (UAT and Production) providing L3 support, maintaining 99%+ uptime through proactive issue tracking and resolution.
    • Validated system resilience through continuous load, integration, and chaos testing cycles — injecting failures to confirm observability, failover, and HA commitments before production impact.
    • Drove platform adoption by implementing an agent and customer referral reward system, increasing user engagement and improving retention metrics.
  • Rokomari.com

    Operations Engineer
    Dec 2024 — Feb 2025
    NGINXModSecurityOpenTelemetry Spring FrameworkPrometheusLokiTempoGrafana
    • Integrated NGINX and ModSecurity WAF with custom rule sets; built symptom-based Grafana dashboards and alerts that surfaced security anomalies before outages, reducing detection and triage time for production traffic incidents.
    • Deployed a unified observability stack with OpenTelemetry, Grafana Loki, Tempo, and Prometheus covering 8+ services (Redis, MongoDB, RabbitMQ, and application tiers), consolidating distributed traces, logs, and metrics into a single pane and reducing average troubleshooting time by over 50%.