Systems Architect

Job ID: 113729
Location: Hampton , Virginia  [On-Site]
Category: App/Dev
Employment Type: Contract
Date Added: 09/30/2026

Apply Now

Fill out the form below to submit your information for this opportunity. Please upload your resume as a doc, pdf, rtf or txt file. Your information will be processed as soon as possible.


 
 
 
 
 
(Word, PDF, RTF, TXT)
⚠Notice: Job application forms are locked until Analytics cookies are accepted for fraud prevention tracking. Please click the cookie settings banner to unlock.
* Required field.

Role Summary

The Site Reliability Engineer supports the resilience, availability, and recoverability of mission-critical cloud-based platforms. This role leads infrastructure-level disaster recovery exercises, validates platform recovery and rebuild procedures, and develops automation to improve operational reliability. The position collaborates with cross-functional engineering teams to support modernization initiatives, strengthen platform resilience, and maintain reliable system performance.

Responsibilities

  • Support the full lifecycle of platform portability and disaster recovery (DR) exercises, including validation of platform rebuild procedures and recovery playbooks.
  • Conduct infrastructure-level recovery exercises to validate that platforms can be rebuilt within established recovery objectives.
  • Verify data completeness, integrity, and accuracy during recovery exercises and document findings and recommended remediation activities.
  • Identify gaps in infrastructure, automation, monitoring, and data recovery processes and collaborate with engineering teams to address identified issues.
  • Design, implement, and support automated Infrastructure as Code (IaC) workflows using technologies such as Terraform and AWS CloudFormation, along with standardized CI/CD pipelines.
  • Manage and optimize Kubernetes clusters and containerized workloads, including provisioning, scaling, and reliability improvements.
  • Develop and maintain observability solutions using CloudWatch, Datadog, or comparable monitoring platforms to support system reliability and proactive incident response.
  • Automate operational processes and develop tools or scripts using languages such as Python or Java to improve efficiency and repeatability.
  • Collaborate with platform engineering, security, application, and data teams to support secure, compliant, and consistent platform operations.
  • Participate in on-call rotations, root-cause analysis, and incident response activities to improve system resilience and operational effectiveness.

Qualifications

  • Bachelor’s degree in a related technical field with a minimum of 5 years of relevant experience in Site Reliability Engineering, DevOps, cloud engineering, infrastructure engineering, or an equivalent combination of education and practical experience.
  • Strong knowledge of AWS services across compute, networking, storage, identity and access management, and serverless technologies.
  • Demonstrated experience with Infrastructure as Code technologies such as Terraform and AWS CloudFormation and infrastructure automation practices.
  • Experience developing and maintaining CI/CD pipelines using GitHub Actions or comparable technologies.
  • Strong knowledge of Kubernetes administration, container orchestration, and containerized deployments.
  • Experience validating disaster recovery processes, performing system recovery or rebuild activities, and conducting data integrity validation.
  • Experience developing and maintaining monitoring and observability capabilities, including dashboards, metrics, logs, and alerts using CloudWatch, Datadog, or comparable platforms.
  • Proficiency with scripting or programming languages such as Python, Java, or C#.
  • Demonstrated ability to troubleshoot complex infrastructure and distributed-system issues, including networking and system reliability challenges.
  • Strong analytical, documentation, problem-solving, and communication skills.
  • Ability to work effectively in a fast-paced environment supporting business-critical systems.
  • Ability to obtain and maintain the required Public Trust determination.

Publishing Pay Range: $38.00 – $42.00 Hourly 
This position is based in office and requires employee to work on-site.