Senior Site Reliability Engineer
Job ID: 113766
Location: Hampton , Virginia [On-Site]
Category: App/Dev
Employment Type: Contract
Date Added: 10/07/2026
Role Summary
A Senior Site Reliability Engineer is responsible for maintaining the reliability, resilience, and recoverability of mission-critical cloud-based platforms. This role involves leading infrastructure disaster recovery drills, validating platform rebuild procedures, and automating deployment workflows. The engineer collaborates across teams to enhance complex environments, support modernization efforts, and ensure operational continuity.
Responsibilities
- Support full lifecycle platform portability and disaster recovery (DR) drill execution, including validation of platform rebuild procedures and DR playbooks
- Execute infrastructure-level drill activities to ensure the platform can be fully rebuilt within the 48-hour recovery target
- Verify end-to-end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations
- Identify exit readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes, and drive corrective actions
- Design, implement, and support automated Infrastructure as Code (IaC) workflows utilizing Terraform, AWS CloudFormation, and CI/CD pipelines
- Manage and optimize Kubernetes clusters and containerized workloads, including provisioning, scaling, and reliability enhancements
- Build and maintain observability solutions using CloudWatch, Datadog, and other monitoring tools for service reliability and proactive incident response
- Develop automation, tooling, and scripts using Python or Java to reduce manual processes and improve operational repeatability
- Collaborate with platform engineering, security, applications, and data teams to ensure secure, compliant, and consistent platform operations
- Participate in on-call rotations, root cause analysis, and incident response activities to strengthen system resilience and operational excellence
Qualifications
- Bachelor’s degree and 8–10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; Master’s degree with 6–8 years; or equivalent practical experience
- Expert knowledge of AWS services across compute, networking, storage, IAM, and serverless components
- Strong experience with Infrastructure as Code (Terraform, CloudFormation) and automation principles
- Experience building CI/CD deployment pipelines and implementing progressive delivery mechanisms using GitHub Actions or similar tools
- Deep understanding of Kubernetes administration, container orchestration, and Docker deployments
- Proven experience validating disaster recovery processes, performing system rebuilds, and conducting data integrity checks
- Experience with monitoring tools like CloudWatch, Datadog, or comparable observability solutions
- Proficiency with programming or scripting languages such as Python, Java, C#, or Go
- Experience debugging complex failure modes, network issues, backpressure, and consistency challenges
- Strong analytical, documentation, and communication skills within technical teams
- Ability to work in a fast-paced environment supporting mission-critical systems
Preferred Qualifications
- AWS DevSecOps Engineer certification (preferred)
- Additional AWS certifications (Solutions Architect, SysOps, Developer) and Kubernetes certifications (CKA, CKAD)
- Familiarity with Zero Trust security models and cloud security best practices
- Experience with GitLab, Jenkins, or similar CI/CD platforms
- Experience working in highly regulated environments, such as healthcare, finance, DHS, DoD, or CMS
- Background supporting federal, defense, or large enterprise programs involving legacy-to-cloud modernization
- Prior involvement in large-scale disaster recovery drills, continuity of operations (COOP), or portability/executable readiness assessments
Publishing Pay Range: $68.00- $72.00 Hourly
This position is based in office and requires employee to work on-site.
