HubviaCareers

SRE Engineer

Atlanta, GA · Contract · Cloud & Infrastructure

About the role

Qualifications

  • Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.
  • Hands-on experience with incident management and 24/7 production support models.
  • Proficiency with monitoring and observability tools such as CloudWatch, Dynatrace, and Quantum Metric.
  • Experience building and maintaining monitoring dashboards.
  • Strong troubleshooting skills across infrastructure, networking, and application layers.
  • Working knowledge of CI/CD pipelines and AWS deployment processes.
  • Experience working with databases and Unix/Linux environments.

Incident Management and Production Support

  • Provide Level 1 and Level 2 support for production incidents across AWS-hosted applications and infrastructure.
  • Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.
  • Escalate code-level defects to development teams with clear diagnostics, supporting logs, and impact assessments.
  • Participate in on-call rotations, major incident bridges, and post-incident reviews.
  • Investigate application defects, configuration issues, and infrastructure anomalies reported through monitoring tools or user incidents.

Monitoring and Operational Health

  • Perform regular health checks across applications, infrastructure, and AWS services.
  • Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and Thousand Eyes.
  • Respond proactively to s related to resource utilization, latency, errors, and availability.
  • Maintain and improve monitoring and observability dashboards.

Mandatory Skills

  • CloudWatch
  • Dynatrace
  • Git
  • Observability
  • Reliability Patterns

Good to Have Skills

Chaos Testing, Shell Scripting

SRE Engineer in Atlanta, GA | Hubvia