About the role

IWU is seeking a Senior Site Reliability Engineer Consultant (AI) (Independent Contractor ‑ 1099 / Contract-to-Hire) for a large-scale technology reorganization project.
ABOUT IWU TECHNOLOGYABOUT THE ROLE: IWU is re-platforming its enterprise on a dual-site, active/active, and on-premises data center foundation and an AI-first application stack. Reliability is a product requirement – not an afterthought bolted on after launch. We are seeking a Senior Site Reliability Engineer Consultant with Sovereign/On-Prem AI Stack Experience – a principal-level individual contributor to define service-level objectives, build the observability platform, and lead incident response for the applications and platform services that serve our 15,000 learners and internal operations. This is not a cloud SRE role. IWU runs production workloads on infrastructure we own and operate in commercial colocation. You will be responsible for application and platform reliability – SLOs, monitoring, alerting, on-call, and post-incident learning – while partnering with Infrastructure on hardware and network availability signals. Application-level performance and error budgets are yours; the data center floor is theirs – and the partnership between the two is what makes uptime real. You will be a hands-on principal: you will translate industry reliability standards into measurable SLOs and deployed tooling, mentor engineers through runbooks and postmortems, and hold the bar for operational excellence across dual sites. Reliability Strategy and SLOsDefine and maintain service-level objectives (SLOs), SLIs, and error budgets for critical platform and application services – availability, latency, throughput, and data freshness – aligned to business impact. Partner with product, application, and platform teams to negotiate realistic targets, document dependencies, and prioritize reliability work against feature delivery. Contribute to the enterprise observability reference architecture – metrics, logs, traces, and synthetic checks – per organizational standards.
Observability Platform: Build AI platform observability in partnership with the Enterprise AI Architect – GPU health (DCGM), inference latency and throughput, model endpoint SLOs, queue depth, and cost/utilization dashboards. Design, deploy, and operate the observability stack – metrics platform, log aggregation, distributed tracing, and dashboards. Implement alerting and escalation – on-call rotations, alert hygiene (actionable alerts only), and runbook linkage. Ensure observability services themselves meet availability and retention targets – dual-site redundancy, capacity planning, and lifecycle management.
Incident Response and Operations: Lead incident response for application and platform outages – triage, communication, mitigation, and resolution – including AI-specific incidents (model degradation, inference failures, data pipeline stalls) where standards apply. Drive post-incident reviews – root-cause analysis, corrective actions, and tracking to completion; feed learnings back into SLOs, runbooks, and architecture. Maintain and improve runbooks, playbooks, and escalation paths – clear ownership, severity definitions, and executive communication templates.
Capacity, Performance, and Resilience: Partner with Infrastructure, Database, and Network teams on capacity signals – trend analysis, saturation forecasting, and proactive scaling before user impact. Design and exercise failure-mode testing – game days, failover drills, and chaos experiments appropriate for dual-site active/active topologies. Support release and deployment reliability – canary patterns, rollback criteria, deployment health checks, and integration with CI/CD pipelines. Analyze performance regressions – APM, tracing, and profiling data to isolate bottlenecks across application, database, and infrastructure layers.
Governance and Collaboration:
  • Align with ITIL incident, problem, and change management – Service
  • Now or equivalent ticketing, CAB participation for high-risk changes, and audit-ready incident records.
  • Partner with Information Security on security incident coordination and observability data handling – retention, access control, and compliance (FERPA, HIPAA, SOC 2) as applicable.
  • Mentor application and platform engineers on reliability patterns – graceful degradation, circuit breakers, idempotency, and building services that are operable by default.
WHAT WE ARE LOOKING FOR: Required
According to Indiana Wesleyan University policy, all employees and contract-to-hire contractors must possess a strong Christian commitment and adhere to the standards outlined in the IWU Community Lifestyle Statement.8+ years of progressive experience in site reliability engineering, production operations, or platform engineering, with 3+ years owning SLOs and on-call for business-critical services. Demonstrated experience building and operating observability at scale – metrics, logs, and traces – on stacks in production. Strong incident management skills – leading Sev-1/Sev-2 response, writing postmortems, and driving permanent fixes rather than repeated firefighting. Experience defining and operating SLO/SLI/error-budget programs – not just uptime percentages, but user-centric reliability targets tied to business outcomes. Hands-on proficiency with Windows, Linux, containers, and Kubernetes operations – pod health, resource limits, HPA, and debugging production clusters. Scripting and automation skills – Python, Go, or Bash – for tooling, alert automation, and toil reduction. Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field – or equivalent demonstrated experience.
Strongly Preferred: Experience with on-premises or colocation production environments – dual-site active/active topologies, and partnering with infrastructure teams on shared incident boundaries. Experience observing AI/ML platform workloads – GPU monitoring, inference SLOs, batch pipeline latency, or model-serving health checks. Familiarity with APM and synthetic monitoring. Experience with Infrastructure as Code – for observability-as-code deployments. Background in a regulated or compliance-sensitive environment (education/FERPA-Title-IV, financial services/GLBA, healthcare/HIPAA, insurance, public sector).Familiarity with ITIL service management and on-call best practices (Google SRE book principles applied pragmatically).Relevant certifications or demonstrated depth equivalent to CKA/CKAD, or similar. How You WorkYou measure before you optimize – SLOs and data drive priorities, not loudest stakeholder or latest outage memory. You design for operability – if on-call cannot diagnose it from the dashboard, the service is not done. You reduce toil relentlessly – automation and self-service beat heroics every time. You mentor by example – your runbooks, SLO docs, and calm incident leadership become the standard others follow.
WHY THIS ROLE?
  • Principal scope, hands-on impact - you drive reliability strategy and the hardest incidents without needing a management title.
  • Greenfield reliability culture. You are building SLO and observability practice alongside a funded dual-site re-platforming – not inheriting a decade of alert spam and no error budgets.
  • AI-native challenge. Inference latency, GPU saturation, and model-serving SLOs are first-class problems here, not edge cases.
  • Clear ownership boundary. Application SLOs are yours; infrastructure hardware is Infrastructure's – clean partnership, real uptime.
  • Mission and meaning. IWU is a century-old Christ-centered university serving roughly 15,000 learners. The reliability you deliver keeps the systems running that prepare them to change the world.
  • PROJECT REMUNERATION:Estimated annual project remuneration is $135,000 - $160,000 for this independent contractor (1099) / Contract-to-Hire role.
  • Contractor Insurance/Bond Requirements
  • General Liability: $1m per occurrence / $2m aggregate
  • Cyber Liability: $2m per occurrence
  • Workers’ Compensation / Employers’ Liability: $500,000; not applicable for approved sole operators
  • Professional Liability / Errors & Omissions: $1m per occurrence / $3m aggregate
  • Equipment & Travel ExpensesIWU will provide required equipment and software. Personal devices and software are generally not permitted for assigned work unless approved by IWU.IWU will pay for required and approved travel and other expenses.
  • Specific policies, rules, and details will be included in the executed contract.
  • Engagement Process
  • Easy-Apply on LinkedInLinkedIn SurveyPhone ScreenVideo Interviews
  • On-Site Interview/Tour
  • Contract Engagement Begins
Compliance & DisclaimersIWU is an equal opportunity employer committed to compliance with Title VII of the Civil Rights Act of 1964, Title IX of the Educational Amendments of 1972 and Section 504 of the Rehabilitation Act of 1973 or other federal, state, or local laws or executive orders except as claimed in a filed religious exemption. Indiana Wesleyan University is managing this search directly. Third-party recruiters, staffing agencies, search firms, subcontracting firms, and candidate marketers may not submit candidates for this role unless they have a current, written agreement with IWU that specifically authorizes recruitment support for this opening. Unsolicited candidate submissions will not create any obligation, fee, commission, placement right, or ownership claim. Any candidate submitted without prior written authorization may be contacted directly by IWU and considered without payment of any recruiting or placement fee. Candidates must apply directly through this LinkedIn posting or the official IWU application process. This role does not authorize third-party representation, bench submissions, or agency-to-agency candidate referrals.

Matching similar jobs

JOB OVERVIEW

Salary

$500,000

Experience level

Manager

Location

Indianapolis, IN

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

yesterday

Tired of running searches?

Rank the roles you'd take once, and matches like these arrive on their own.

CREATE PROFILE