About the Role
Our client is a top-tier strategy and management consulting firm, the kind that competes with McKinsey, BCG and Bain for C-suite work. They are building an AI-native development platform that standardizes how code is generated, verified, secured, deployed, and monitored, so AI-assisted product development is safe and repeatable across its portfolio.
The Principal Platform Engineer sits at the intersection of infrastructure, delivery engineering, security, and developer experience. You will be accountable for the substrate every product - infrastructure as code, delivery pipelines and their gates, the runtime, identity and network, security and compliance controls, telemetry and cost, and the packaged development environment that carries our AI-assisted engineering harness.
This is a hands-on engineering and technical leadership position. You will move fluidly between building production infrastructure, architecting multi-tenant isolation, designing pipeline and policy enforcement, leading strategic architecture decisions, and mentoring engineers. Your success is measured by what every product inherits: a compounding system that outlasts your direct involvement. Travel is part of this position, but frequency may vary based on client, team, and individual circumstances.
What You'll Do:
Platform Architecture & Infrastructure Engineering
● Drive the platform's architecture and runtime direction, including convergence onto one container-orchestrated runtime, tenancy-grade isolation between client engagements, and the contracts every product inherits.
● Design composable infrastructure as code at fleet scale, with clear input contracts and state boundaries, published as versioned templates and images for every core component.
● Make the environment lifecycle a one-command operation: created by merging a pull request, destroyed by deleting it, with automated proof nothing is left behind.
● Design identity and network as a federated, credential-free system — one identity per engagement, just-in-time audited privileged access, separate internal and external edges.
● Ship telemetry and cost visibility out of the box for every service: standard alerting, service-level objectives on live data, fault injection, and per-engagement cost attribution including AI spend.
Delivery Pipelines, Security & Compliance Engineering
● Drive the shared delivery pipelines every product runs through, and the pluggable gate framework that quality, security, and AI-governance checks plug into with block-on-fail enforcement.
● Implement security and compliance as code: release-blocking scanning on every repository, policy and admission control, signed and attested artifacts, and an audit trail and software bill of materials, including AI components, on every build as SOC2 / ISO 27001 evidence.
● Package the AI-assisted development environment as a versioned, first-class deliverable, and enforce the sandbox boundary for coding agents: default-deny egress, short-lived credentials, isolated shared runners.
● Run the platform as a product measured by developer outcomes — setup time, time to first commit, pipeline duration, gate false-positive rate, adoption — pruning or redesigning controls when governance burden outgrows benefit, and proving every golden path by shipping through it yourself. Technical Leadership & Stakeholder Influence
● Lead the platform's strategic architecture decisions — multi-tenancy, cloud posture, identity provider, network edge — as evidence-based decision documents with real trade-offs, and own the security and reliability evidence the program reports against.
● Serve as a go-to technical authority on platform architecture and delivery engineering, consulted by senior stakeholders at the design and strategy stages, translating trade-offs and platform risk into actionable guidance, and mentoring engineers and vendor partners with observable improvement in design quality and production readiness.
What You'll Need:
● 15+ years for the Staff level and 10+ years for the Principal level - software engineering, platform or infrastructure engineering, or a closely related technical field.
● Deep, hands-on expertise in infrastructure as code at fleet scale (Terraform or OpenTofu with Terragrunt or equivalent): module architecture, remote state, multi-environment topology.
● Demonstrated experience operating multi-tenant Kubernetes platforms in production: GitOps, network and admission policy, workload identity, tenant isolation.
● Proven track record building CI/CD as a shared platform (GitHub Actions or equivalent), with reusable workflow contracts, fleet-scale branch protection, and enforced quality and security gates.
● Strong hands-on command of cloud identity and network security on Azure/AWS/GCP or a comparable hyperscaler: workload identity federation, RBAC, private networking, WAF, DNS.
● Substantive experience with security and compliance as code: scanning toolchains (Snyk, Checkov, Trivy), policy engines (OPA, Kyverno, Azure Policy), supply chain controls (signing, provenance, SBOM), and SOC2 / ISO 27001 control evidence.
● Production observability with OpenTelemetry or equivalent: collector topology, SLOs and error budgets, alert routing, incident management, fault injection.
● Has built and operated production application software as a software engineer, in at least one of Go, Python, or TypeScript, and is fluent in testing strategy, code review, trunk-based delivery, and release management.
● Excellent written and verbal communication skills in English; ability to translate complex technical topics to diverse audiences, including executive stakeholders.
● Experience with Agile methodologies and cross-functional product team collaboration.
Preferred / Additional Qualifications
● Experience building platforms for consulting, advisory, or professional services, including client-isolation and data-residency requirements.
● Familiarity with virtual-cluster and control-plane tooling (vcluster, Crossplane) and developer portals (Backstage, Port).
● Experience operating a platform used by AI coding assistants (Claude Code, GitHub Copilot or similar), including sandboxing of agent tooling.
● Advanced certifications in cloud architecture, Kubernetes, or security (e.g., Azure Solutions Architect, CKA/CKS, CISSP). ● FinOps experience: cost allocation, showback, and cloud cost optimization.
● Demonstrated enthusiasm for developer education — writing internal guides, running workshops, or building internal tooling communities.