About the role
What will you do at Akamai Technologies?
Do you want to shape reliability practices for a new AI inference platform?
Are you a senior technical leader who drives solutions across teams?
Join the Akamai Inference Cloud Team!
The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We
design, implement, deploy and operate AI platforms that enable customers to run
inference models and developers to create AI applications.
Partner with the best
In this role, you'll lead reliability workstreams for Akamai's serverless
inference platform, design SRE tooling and automation, and drive technical
decisions. Opportunities exist to mentor other SREs, influence architecture
decisions with product engineering teams, and shape SRE practices for AI
inference workloads and GPU infrastructure at scale.
As a Senior II Site Reliability Engineer, you will be responsible for:
* Taking ownership of observability strategy for the serverless inference
platform, designing telemetry, dashboards, and alerts, defining SLO/SLI
frameworks, and driving improvements when targets are missed
* Building production-grade automation and tooling that reduces operational
toil, improves incident response, and sets patterns that other SREs adopt
* Owning incident management integration for inference workloads, designing
frameworks, leading incident response during on-call rotations, and driving
systemic improvements from post-mortems
* Defining and implementing deployment safety practices including progressive
rollouts, canary analysis, and rollback automation, establishing standards
for the team
* Partnering with product engineering teams to influence architecture
decisions, ensure operational readiness, and represent the SRE perspective in
design reviews
* Mentoring Senior and mid-level SREs through code reviews, design discussions,
and hands-on problem-solving
Do what you love
To be successful in this role you will:
* 8+ years of experience in SRE, infrastructure engineering, or platform
engineering, working with large-scale distributed systems
* Possess a proven track record of defining SLO/SLI frameworks, building
observability platforms, and running incident management processes at scale
* Have extensive Kubernetes and containerization experience at scale, including
autoscaling, resource scheduling, and container orchestration for
compute-intensive workloads
* Have experience building automation and tooling in Python or Go, with
familiarity in CI/CD pipelines, deployment safety, and infrastructure-as-code
* Possess the ability to lead technical initiatives across teams, mentor other
engineers, and drive complex reliability problems to resolution independently
* Have experience with or exposure to AI/ML infrastructure, model serving, or
GPU workloads
About us
At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your
favorite series, or managing your savings, we're the engine behind the scenes.
We provide the world's most distributed platform from Cloud to Edge to help the
giants of the digital world work faster and stay more secure, making the
internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's
most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because
behind every great digital experience is a massive hidden challenge. And we're
the ones who solve it. When millions of people hit play or pay, Akamai ensures
it just works.
Benefits
at Akamai: We support your health, well-being, finances, and life
beyond work. See our benefits.
[https://www.akamai.com/careers/working-at-akamai? utm_source=akamai&utm_medium=social&utm_campaign=benefits]
FlexBase adapts to your job's needs
Akamai's FlexBase program is yet another way we show our commitment to providing
employees with an exceptional workplace experience. It's not about telling
employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home,
in an office, or a combination of both.
Connect with us on social and see what life at Akamai is like!
[https://app.role-mapper.com/images/twitter.png]https://twitter.com/Akamai? ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Eauthor
[https://app.role-mapper.com/images/facebook.png]https://www.facebook.com/akamaicareers/
[https://app.role-mapper.com/images/linkedin.png]https://www.linkedin.com/company/akamai-technologies/
[https://app.role-mapper.com/images/instagram.png]https://www.instagram.com/akamaicareers/
[https://app.role-mapper.com/images/blog.png]https://www.akamai.com/blog?
[https://assets.role-mapper.com/images/youtube_new.png]https://www.youtube.com/playlist? list=PLYo5nh8xQFpnnPl17sQRURR9ugKL-KMRO
Compensation
Akamai is committed to fair and equitable compensation practices. For US based
candidates only - the base salary for this position ranges from $146,400 -
$263,600/year; a candidate’s salary is determined by various factors including,
but not limited to, relevant work experience, skills, certifications and
location.
Compensation
for candidates outside the US will vary. The compensation
package may also include incentive compensation opportunities in the form of
annual bonus or incentives, equity awards and an Employee Stock Purchase Plan
(ESPP). Akamai provides industry-leading benefits including healthcare, 401K
savings plan, company holidays, vacation (in the form of PTO), sick time, family
friendly benefits including parental leave and an employee assistance program
including a focus on mental and financial wellness; Eligibility requirements
apply.
Which skills does this role require?
Make your next move
Build a shortlist and prepare
Identify the requirements you can demonstrate, then choose examples from your work to discuss with the hiring team.
- Build a focused shortlist before you applyCompare role requirements with your experience and give each application a clear reason.
- Platform engineering jobsCompare current openings and review what to look for in this role.
- Practice explaining your experience in an interviewRehearse your answers before meeting the hiring team.
Other roles to compare
Review the responsibilities and requirements before adding an opening to your shortlist.
Machine Learning Infrastructure Engineer
Bright Vision Technologies · Hillsboro, Oregon, United States
Senior AI Infrastructure Software Engineer - DGX Cloud
NVIDIA · Redmond, Washington, United States
Security Engineer - Infrastructure Security
Figure · San Jose, California, United States
Data and ML Infrastructure Engineer
HavocAI · United States
Senior Machine Learning Engineer - ML Training Infrastructure
General Motors · Sunnyvale, California, United States
Software Engineer, Infrastructure Services (Data Plane)
Apple · California, United States
Role information can change. Confirm current details on the original application page.
