RVC
JobsFor Employers
16 jobsSort: Relevance
5d 1h ago

NOC Technician

SpaceXAI·General Helpdesk/IT Support
📡Remote In-Country
|United States of America
monitoring
6d 4h ago

Hotfix Developer

Paystone·General Helpdesk/IT Support · Payment Processing and Software
📡Fully Remote
dockergitgraylogphp+2
10d 14h ago

Advanced Technical Support Specialist

SupportYourApp·General Helpdesk/IT Support · Support Services
📡Fully Remote
command_linelinuxnetworkingwindows
11d 1h ago

Purview and Copilot Expert

AlmavivA de Belgique·General Helpdesk/IT Support
📡Fully Remote
compliancemonitoringrisk_managementmicrosoft_sharepoint
11d 19h ago

IT Operations Specialist

Devoted Studios·General Helpdesk/IT Support · Game Development
📡Fully Remote
jiraversion_control_systems
12d 3h ago

NOC Administrator

True Zero Technologies·General Helpdesk/IT Support · Engineering & Architecture
📡Fully Remote
ciscoitiltroubleshooting
12d 20h ago

Technical Support Specialist

WOW 24-7·General Helpdesk/IT Support · E-commerce / Marketplace
📡Fully Remote
$1000 - $1400 USD
mssqlnetworkswindows_serverzendesk
13d 10h ago

Technical Support Specialist

General Helpdesk/IT Support · Fintech
📡Fully Remote
apigrafanajsonpostman+1
16h 1m ago

IT Support Engineer

Runway·General Helpdesk/IT Support · Artificial intelligence software
📡Fully Remote
|London
£65k - £90k GBP
awsbashci/cddhcp+10
7d 20h ago

IT Operations Manager

Echo Base·General Helpdesk/IT Support · Investment
📡Fully Remote
awsbashcloudflarecms+3
8d 12h ago

Senior Corporate IT Engineer

Yuno·General Helpdesk/IT Support · Technology
📡Fully Remote
10d 8h ago

Network Service Delivery Manager

New Era Technology·General Helpdesk/IT Support
📡Fully Remote
$90k - $94.5k USD
saasteam_management
13d 2h ago

Senior System Administrator

General Helpdesk/IT Support · IT Services
📡Fully Remote
active_directoryansiblebashdns+8
13d 3h ago

Principal Technical Support Engineer

Mitek·General Helpdesk/IT Support · Digital Identity Authentication
📡Fully Remote
€37k - €56k EUR
ecseksgithubgrpc+10
13d 4h ago

Technical Support Engineer

Truv·General Helpdesk/IT Support · Financial Services
📡Fully Remote
$45k - $70k USD
analyticscommunication_skillszendesk
1d 20h ago

Technical Support Specialist

Innovation Systems·General Helpdesk/IT Support · Green industry
📡Fully Remote
$1000 - $1400 USD
bug_reportingcrmjirazendesk

NOC Technician

SpaceXAI | General Helpdesk/IT Support
Remote In-Country
Full Time
United States of America

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a NOC Technician, you are the eyes and the voice of the site — never the hands. You staff the Network / Campus Operations Center and continuously observe site health signals across xAI campuses. You detect and verify campus-impacting events, assemble the right responders, run incident communications leadership can trust, and drive every major incident to a completed report and a tracked corrective project. You work with Site Reliability Engineering, SiteOps, Facilities, Hardware Failure Analysis, SWE Platforms, and vendors — escalating correctly the first time and maintaining the institutional memory across shifts and sites.
One sentence: watch the campus, run the bridge, leave the wrench work and deep root cause to the teams that own them.

RESPONSIBILITIES:
Continuous monitoring (the watch)
• Staff the console per shift schedule to sustain 24/7 coverage (coverage posture: 2 on console per site)
• Watch the designated signal surface: cluster health dashboards, node availability, network health, facility trend panels (power/cooling), storage alarms, and threshold breaches as defined by SRE monitoring standards
• Acknowledge every page/alert within the SLA; classify it (actionable / known / noise) and log the disposition; feed noise patterns back to SRE for suppression or redesign
• Maintain a live picture of ongoing maintenance, planned work, and degraded-but-accepted states so real anomalies stand out

Detection, triage & escalation
• Detect → verify → escalate within defined time budgets; verification is signal-level (is it real, what's the blast radius), not deep diagnosis
• Operate the escalation matrix: NOC → on-call SRE → domain owners (SiteOps, Facilities, Network, Storage, HW FA, vendors); page correctly the first time
• Recommend incident declaration and severity to the on-call SRE; declare directly per runbook when thresholds are unambiguous

Incident communications & coordination
• Open and run the bridge; get the right people on within the time-to-bridge SLA
• Own stakeholder communications: first update within the SLA, then a fixed cadence until resolution
• Maintain the incident timeline in real time — timestamps, actions, decisions, engagements
• Track who owns what during the incident and call out stalls

First-pass RCA framing & closure
• Produce initial framing for major site outages: what happened, when it started, what's impacted (halls/racks/services), what changed recently, who is engaged
• Hand framing to SRE / Hardware FA for depth — the NOC does not publish root cause
• Write major-incident reports; open corrective projects in Linear with named owners and track them to closure ("filed" is not "done")

Shift operations, runbooks & improvement
• Run structured shift handoffs and keep durable shift logs; maintain cross-site awareness
• Own and continuously improve NOC runbooks: escalation matrix, comms templates, severity ladders, per-signal response procedures
• Participate in game days run by SRE; every incident where the runbook was wrong or missing produces a runbook change before the incident closes

Explicitly not this role
• Wrench work: swaps, reseats, physical recovery (SiteOps)
• Power / cooling / building plant operation (Facilities)
• Deep hardware root-cause analysis or vendor CAPA (Hardware Failure Analysis)
• Monitoring architecture, alert design, or technical SEV command (Site SRE)
• Building or operating reliability tooling such as SRT, turnback, or dashboards (SWE Platforms)

BASIC QUALIFICATIONS:
• High school diploma or equivalency certificate
• 1+ year of professional experience in a Network Operations Center (NOC), Security Operations Center (SOC), mission-control / dispatch, data center operations watch, or equivalent 24/7 monitoring and incident-communications role
• Demonstrated written and verbal communication skills under time pressure (stakeholder updates, handoffs, timelines)

PREFERRED SKILLS AND EXPERIENCE:
• Calm under pressure; excellent written and verbal communications — leadership should be able to trust your incident updates verbatim
• Pattern recognition across domains; multi-domain curiosity (compute, network, storage, power/cooling signals)
• Experience following and improving process: runbooks, escalation matrices, shift handoffs, post-incident follow-through
• Prior NOC, SOC, or critical-environment operations experience in a datacenter or hyperscale infrastructure environment
• Familiarity with reading operational dashboards, acknowledging/classifying alerts, and coordinating across on-site technicians, facilities, and engineering on-call
• Comfort with ticketing / project tracking systems (e.g. Linear, Jira) for opening and chasing corrective work to closure
• Industry certifications a plus (Network+, Security+, ITIL, or similar) — not a substitute for judgment and communications quality
• Basic familiarity with datacenter topology (racks, fabric, OOB) and how facility events affect compute availability — enough to triage and escalate correctly, not to deep-diagnose
• Bachelor's degree in IT, Computer Science, Cybersecurity, or STEM discipline preferred but not required

ADDITIONAL REQUIREMENTS:
• Must be available for on-shift rotations supporting 24/7/365 console coverage
• Shift structure (e.g. 12-hour rotations) to be confirmed; nights, weekends, and holidays are part of the role
• Must be able to work extended hours during major incidents as needed

SUCCESS LOOKS LIKE:
• Coverage attainment / shift fill rate vs plan
• Time-to-bridge for major incidents; first-update and cadence SLA attainment
• Page accuracy / escalation correctness
• % of major incidents with a complete timeline and follow-up projects tracked to done
• Not measured by: raw page counts, or heroics without a paired prevention item

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Required Skills

monitoring

Required Languages

🇬🇧 English

Key competency: General Helpdesk/IT Support
Roles
PythonJavaReactTypeScriptNode.jsGoRustDevOpsData scienceProductDesign
Remote
United StatesUnited KingdomCanadaGermanyPolandSpainNetherlandsPortugal
Work type
Fully remoteRemote in-countryHybridSeniorMid-levelJuniorAll jobs
R© 2026 ReVacancybuild e3cb608f
AboutContactPrivacyCookiesRefundsTerms & ConditionsFor Employers