RVC
JobsFor Employers
  1. Home
  2. /
  3. Jobs
  4. /
  5. Member of Technical Staff, Site Reliability Engineer

Member of Technical Staff, Site Reliability Engineer

Inferact | Artificial Intelligence
1 hour 35 minutes ago
Remote In-Country
Full Time
United States of America, San Francisco
$200k - $400k USD/yr
Source not independently verified — confirm the company domain before applying.

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes.

You'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar.

  • Strong experience operating production systems with meaningful traffic, user impact, or infrastructure criticality.

  • Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.

  • Experience live-fighting major production incidents, including mitigation, root cause analysis, escalation, and follow-through on prevention work.

  • Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.

  • Ability to design operationally simple systems and identify likely failure modes before launch.

  • Strong programming or scripting ability in Python, Go, Bash, or similar for automation, tooling, and reliability improvements.

Preferred qualifications:

  • Experience supporting ML infrastructure, inference systems, GPU workloads, Kubernetes-based platforms, or high-scale backend services.

  • Experience building or improving observability systems using metrics, logs, traces, dashboards, alerts, and runbooks.

  • Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD systems, or production deployment platforms.

  • Experience driving incident review culture, post-mortem processes, reliability reviews, and prevention-oriented engineering work.

  • Ability to partner with engineering teams to improve service design, release safety, capacity planning, and operational readiness.

Bonus points if you have:

  • Owned reliability for high-throughput, latency-sensitive, or mission-critical production systems.

  • Supported AI inference, model serving, GPU clusters, ML platforms, or distributed serving infrastructure.

  • Built automation that reduced toil, improved recovery time, or prevented repeat incidents.

  • Led incident response for severe outages with clear communication across engineering and leadership.

  • Created practical SLOs, dashboards, alerts, runbooks, or release gates that improved production reliability.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: We offers generous health, dental, and vision benefits as well as 401(k) company match.

Required Skills

bashci/cddistributed_systemsdockergolangkuberneteslinuxnetworkingpythonterraform

Required Languages

🇬🇧 English

11 jobs
Sort:
1h 35m ago

Member of Technical Staff, Site Reliability Engineer

Inferact·Artificial Intelligence
📡Remote In-Country
|San Francisco, United States of America
$200k - $400k USD/yr
bashci/cddistributed_systemsdocker+6
Similar jobs
42m ago

Site Reliability Engineer

Xsolla·Gaming commerce
🏢On-site
|Baku, Azerbaijan
bashci/cddatadoggcp+15
1h 36m ago

Senior Software Engineer

Scale AI·Artificial intelligence technology
🏢On-site
|Doha, Qatar|Relocation
awsazureci/cdgcp+6
2h 34m ago

Site Reliability Engineer Intern

ShopBack·E-commerce and payments
🏢On-site
|Singapore, Singapore
ai_toolsalgorithmsautomationaws+17
7h 7m ago

Cloud Software Engineer

ClickHouse·Cloud Software
📡Remote In-Country
|Canada
awsazureclickhousedistributed_systems+9
8h 17m ago

Senior Production Engineer

Lightspark·Financial infrastructure
📡Remote In-Country
|Los Angeles, United States of America|Relocation
$200k - $238k USD/yr
awsdebuggingdevopsinfrastructure_as_code+5
13h 41m ago

Staff Site Reliability Engineer

Garner Health·Healthcare
📡Remote In-Country
|New York City
$241k - $270k USD/yr
awsdatadoggitlabgolang+9
16h 49m ago

Head of Platform & Solutions Operations (f/m/d)

B. Braun·Medical Equipment and Pharmaceutical Manufacturing
🏢On-site
|Melsungen, Germany
gitopsiso_27001kubernetes
22h 39m ago

Observability Engineer

Bitdeer·AI and cloud infrastructure
📡Remote In-Country
|San Jose, United States of America
golangkubernetesopentelemetryprometheus+2
1d 6h ago

Site Reliability Engineer

Anduril Industries·Defense technology
🏢On-site
|Waltham, United States of America
$166k - $220k USD/yr
bashdevopslinuxpython+2
1d 11h ago

NOC Analyst

IDT Corporation·Telecommunications
📡Fully Remote
|Brasilia
awsazurebgpgrafana+9
Roles
PythonJavaReactTypeScriptNode.jsGoRustDevOpsData scienceProductDesign
Remote
United StatesUnited KingdomCanadaGermanyPolandSpainNetherlandsPortugal
Work type
Fully remoteRemote in-countryHybridSeniorMid-levelJuniorAll jobs
R© 2026 ReVacancybuild bc959c70
AboutContactPrivacyCookiesRefundsTerms & ConditionsFor Employers