Lead DevOps Engineer (Remote - US)
The Work: ICF is seeking a Lead DevOps Engineer to own and evolve cloud platforms that support modern enterprise applications. The role will operate multi-account AWS environments, strengthen infrastructure automation and GitOps delivery, and partner with application, security, database, and enterprise cloud teams to provide secure, observable, and recoverable services.
Job Location: This position requires that the job be performed in the United States. If you accept this position, you should note that ICF monitors employee work locations, blocks access from foreign locations and foreign IP addresses and prohibits personal VPN connections.
What You Will Do:
- Own the architecture and day-to-day operation of multi-account AWS environments across development, QA, staging, and production, including VPC, IAM, EKS, Elastic Load Balancing, CloudFront, Route 53, ACM, S3, and ECR.
- Develop and maintain reusable Terraform modules and Terragrunt environment configurations, including remote state, stack dependencies, drift detection, change planning, approvals, imports, and recovery procedures.
- Administer Amazon EKS Auto Mode and Kubernetes resources using Helm, Argo CD, and ApplicationSets; manage upgrades, access, capacity, ingress, secrets, and workload reliability.
- Support GitOps controllers and Kubernetes infrastructure automation, including AWS Controllers for Kubernetes and Kubernetes Resource Orchestrator, and resolve reconciliation, ownership, and lifecycle failures.
- Build and maintain GitHub Actions workflows, AWS OIDC authentication, self-hosted runners, container-image pipelines, environment promotion, deployment approvals, and rollback procedures.
- Operate Microsoft SQL Server on Amazon RDS, including parameter and option groups, backups, snapshots, point-in-time recovery, restore testing, Performance Insights, CloudWatch monitoring, and AWS DMS migration workflows.
- Implement security and audit controls using IAM, KMS, Secrets Manager, AWS Config, CloudTrail, GuardDuty, Security Hub, Inspector, WAF, security groups, and centralized logging.
- Build and maintain observability using Prometheus, Grafana, Amazon CloudWatch, OpenTelemetry and AWS Distro for OpenTelemetry, with actionable dashboards, metrics, logs, traces, alerts, and service-level objectives.
- Partner with developers to review and troubleshoot Java and Spring Boot services, including application startup, JVM performance, API behavior, configuration, database connectivity, and container or Kubernetes deployment failures; use logs, metrics, and traces to distinguish application defects from infrastructure issues.
- Integrate application and supply-chain security into delivery pipelines, including SAST with SonarQube, DAST, software composition analysis, container-image scanning, dependency checks, and release quality gates.
- Lead incident response, root-cause analysis, disaster-recovery exercises, infrastructure upgrades, cost optimization, architecture documentation, operational runbooks, and knowledge transfer.
What You Will Bring With You:
- Bachelor's degree in computer science, information technology, engineering, or a related field, or equivalent professional experience.
- 8+ years of DevOps, cloud infrastructure, platform engineering, or site reliability experience, including 5+ years of experience operating production workloads on AWS.
- 5+ years of experience with Terraform and infrastructure as code (IaC), including reusable modules, remote state, imports, drift reconciliation, and automated plan and apply workflows
- 4+ years of experience administering production Kubernetes environments, including Amazon EKS, Helm, ingress, RBAC, secrets, upgrades, observability, and GitOps delivery with Argo CD or a comparable platform.
- 3+ years of experience building and supporting CI/CD pipelines with GitHub Actions or a comparable platform, plus strong Linux, scripting, AWS networking, IAM, troubleshooting, and incident-response skills.
- Must be a US Citizen or Permanent Resident per contract requirements.
What We Would Like You To Bring With You:
- Experience taking ownership of and improving an existing production cloud platform, including its architecture, automation, reliability, security, and operational practices.
- Hands-on experience with Terragrunt, EKS Auto Mode, AWS Pod Identity, Argo CD ApplicationSets, AWS Controllers for Kubernetes, or Kubernetes Resource Orchestrator.
- Experience operating Microsoft SQL Server on Amazon RDS and supporting AWS DMS, backup and restore, disaster recovery, and database performance monitoring.
- Experience with CloudFront, Route 53, ACM, Cognito, SQS, SNS, SES, Secrets Manager, S3, ECR, and cross-account AWS access patterns.
- Experience implementing observability with Prometheus, Grafana, CloudWatch, OpenTelemetry, centralized logging, alerting, and service-level objectives.
- Experience implementing SAST with SonarQube, DAST, software composition analysis, container scanning, SBOMs, and policy-based release gates.
- Hands-on experience reviewing and troubleshooting Java and Spring Boot applications, including REST or GraphQL APIs, Maven, JVM diagnostics, database connection pools, and containerized deployments.
Bot and Third-Party Applications
Working at ICF
ICF is a global advisory and technology services provider, but we’re not your typical consultants. We combine unmatched expertise with cutting-edge technology to help clients solve their most complex challenges, navigate change, and shape the future.We can only solve the world's toughest challenges by building a workplace that allows everyone to thrive. We are an equal opportunity employer. Together, our employees are empowered to share their expertise and collaborate with others to achieve personal and professional goals. For more information, please read our EEO policy.
We will consider for employment qualified applicants with arrest and conviction records.
Read more about workplace discrimination rights or our benefit offerings which are included in the Transparency in (Benefits) Coverage Act.
At ICF, we are committed to ensuring a fair interview process for all candidates based on their own skills and knowledge. As part of this commitment, the use of artificial intelligence (AI) tools to generate or assist with responses during interviews (whether in-person or virtual) is not permitted. This policy is in place to maintain the integrity and authenticity of the interview process.
Pay Range - There are multiple factors that are considered in determining final pay for a position, including, but not limited to, relevant work experience, skills, certifications and competencies that align to the specified role, geographic location, education and certifications as well as contract provisions regarding labor categories that are specific to the position.
The pay range for this position based on full-time employment is:
$131,256.00 - $223,134.00 Nationwide Remote Office (US99)Required Skills
Required Languages
🇬🇧 English