Senior Site Reliability Engineer

Job ID: 113766
Location: Hampton , Virginia  [On-Site]
Category: App/Dev
Employment Type: Contract
Date Added: 10/07/2026

Apply Now

Fill out the form below to submit your information for this opportunity. Please upload your resume as a doc, pdf, rtf or txt file. Your information will be processed as soon as possible.


 
 
 
 
 
(Word, PDF, RTF, TXT)
⚠Notice: Job application forms are locked until Analytics cookies are accepted for fraud prevention tracking. Please click the cookie settings banner to unlock.
* Required field.

Role Summary
A Senior Site Reliability Engineer is responsible for maintaining the reliability, resilience, and recoverability of mission-critical cloud-based platforms. This role involves leading infrastructure disaster recovery drills, validating platform rebuild procedures, and automating deployment workflows. The engineer collaborates across teams to enhance complex environments, support modernization efforts, and ensure operational continuity.

Responsibilities

  • Support full lifecycle platform portability and disaster recovery (DR) drill execution, including validation of platform rebuild procedures and DR playbooks
  • Execute infrastructure-level drill activities to ensure the platform can be fully rebuilt within the 48-hour recovery target
  • Verify end-to-end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations
  • Identify exit readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes, and drive corrective actions
  • Design, implement, and support automated Infrastructure as Code (IaC) workflows utilizing Terraform, AWS CloudFormation, and CI/CD pipelines
  • Manage and optimize Kubernetes clusters and containerized workloads, including provisioning, scaling, and reliability enhancements
  • Build and maintain observability solutions using CloudWatch, Datadog, and other monitoring tools for service reliability and proactive incident response
  • Develop automation, tooling, and scripts using Python or Java to reduce manual processes and improve operational repeatability
  • Collaborate with platform engineering, security, applications, and data teams to ensure secure, compliant, and consistent platform operations
  • Participate in on-call rotations, root cause analysis, and incident response activities to strengthen system resilience and operational excellence

Qualifications

  • Bachelor’s degree and 8–10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; Master’s degree with 6–8 years; or equivalent practical experience
  • Expert knowledge of AWS services across compute, networking, storage, IAM, and serverless components
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and automation principles
  • Experience building CI/CD deployment pipelines and implementing progressive delivery mechanisms using GitHub Actions or similar tools
  • Deep understanding of Kubernetes administration, container orchestration, and Docker deployments
  • Proven experience validating disaster recovery processes, performing system rebuilds, and conducting data integrity checks
  • Experience with monitoring tools like CloudWatch, Datadog, or comparable observability solutions
  • Proficiency with programming or scripting languages such as Python, Java, C#, or Go
  • Experience debugging complex failure modes, network issues, backpressure, and consistency challenges
  • Strong analytical, documentation, and communication skills within technical teams
  • Ability to work in a fast-paced environment supporting mission-critical systems

Preferred Qualifications

  • AWS DevSecOps Engineer certification (preferred)
  • Additional AWS certifications (Solutions Architect, SysOps, Developer) and Kubernetes certifications (CKA, CKAD)
  • Familiarity with Zero Trust security models and cloud security best practices
  • Experience with GitLab, Jenkins, or similar CI/CD platforms
  • Experience working in highly regulated environments, such as healthcare, finance, DHS, DoD, or CMS
  • Background supporting federal, defense, or large enterprise programs involving legacy-to-cloud modernization
  • Prior involvement in large-scale disaster recovery drills, continuity of operations (COOP), or portability/executable readiness assessments

Publishing Pay Range: $68.00- $72.00 Hourly

This position is based in office and requires employee to work on-site.