About the role
What will you do at Citus Data?
Incident triage and first-line response: Provide on-call coverage for incoming incidents across CDI services. Perform initial investigation, severity assessment, and routing to owning engineering teams. Agentic triage system development: Build and extend AI-driven agents that ingest ICM alerts, correlate with recent deployments and feature flag rollouts, check known-issue databases, and produce initial assessments with suggested severity and owning team.
TSG and known-issue matching: Develop automation that matches incoming incidents to relevant Troubleshooting Guides (TSGs) and known issues across Fabric and Power Platform — reducing investigation time and enabling faster resolution. 4+ years of software engineering experience in site reliability, Live site operations, or incident management for cloud services. Good programming skills in one or more of: C#, PowerShell, Python, KQL/Kusto.
Experience with incident management systems and workflows (ICM, PagerDuty, ServiceNow, or similar). Experience with monitoring, alerting, and observability systems (Kusto, Geneva, Grafana, or similar). Ability to work in an on-call rotation across time zones in a geographically distributed team.
Experience interface with engineers, leadership, support, and customers. Experience building AI/ML-driven automation, agents, or intelligent workflows (e.g., using LLMs, Copilot extensibility, MCP servers, or agentic frameworks). Familiarity with Live site ecosystem management (including log traversal, incident management, telemetry analysis, etc.)
Experience with Azure, Power BI, and Fabric services. Experience with Troubleshooting Guide (TSG) authoring and incident pattern analysis. Understanding of SLA management, customer communications, and escalation workflows for cloud services.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Data and ML Infrastructure Engineer
HavocAI · United States
Senior Machine Learning Engineer - ML Training Infrastructure
General Motors · Sunnyvale, California, United States
Neural Data Infrastructure Engineer
Blackrock Neurotech · Salt Lake City, Utah, United States
Role information can change. Prepin can help you prepare, but does not submit an application for this role.