logo

AI Operations Manager

Admiral Talent > Jobs > AI Operations Manager

About the Opportunity

Artificial intelligence is rapidly moving from experimentation into the operating core of the enterprise.

Organizations are no longer managing a handful of isolated machine-learning models.

They are operating generative AI applications, enterprise copilots, retrieval systems, foundation-model APIs, AI agents, intelligent workflows, internal AI platforms, specialized models, vector databases, evaluation frameworks, and third-party AI services—often across multiple business functions simultaneously.

That transition creates an entirely new operational challenge.

An AI system can technically be online and still be failing.

The infrastructure may be healthy while response quality deteriorates.

A model provider may change behavior without a traditional application deployment.

An agent may successfully complete thousands of tasks before an edge case exposes an unexpected control weakness.

Token consumption can increase rapidly.

Retrieval quality can deteriorate as enterprise knowledge changes.

Latency can increase while traditional infrastructure metrics remain within acceptable thresholds.

And a seemingly small prompt, model, policy, data, or configuration change can materially alter how an AI application behaves.

We are seeking an experienced AI Operations Manager to build the operational discipline required to run enterprise AI reliably, securely, responsibly, and economically at scale.

This position will sit at the intersection of Artificial Intelligence, Machine Learning Engineering, Platform Engineering, Cloud Infrastructure, Data, Cybersecurity, Risk, FinOps, Product, and Business Operations.

The AI Operations Manager will help establish how production AI systems are deployed, monitored, evaluated, governed, supported, optimized, and continuously improved.

The mandate is substantially broader than traditional infrastructure operations.

You will need to understand both system health and AI behavior.

That means monitoring conventional signals such as uptime, latency, throughput, errors, capacity, and infrastructure utilization while also establishing visibility into AI-specific signals such as response quality, hallucination rates, retrieval performance, model drift, token utilization, inference cost, safety violations, evaluation results, agent completion rates, escalation patterns, and user adoption.

Generative AI will be a major component of the role.

The environment may include enterprise use of platforms and models from leading AI providers alongside open-source models, internally developed solutions, retrieval-augmented generation, AI gateways, vector databases, orchestration frameworks, and emerging agentic systems.

The AI Operations Manager will help ensure these capabilities can move safely from prototype → controlled deployment → production → enterprise scale.

AI governance will be equally important.

This leader will partner with Security, Legal, Privacy, Risk, Compliance, Data Governance, and business stakeholders to establish practical controls around model access, sensitive information, third-party AI services, responsible AI, human oversight, logging, evaluation, change management, and incident response.

The goal is not to create governance that prevents innovation.

The goal is to create an operating environment where the organization can innovate faster because appropriate controls already exist.

Economics will also matter.

As enterprise AI adoption grows, inference, model API, GPU, storage, retrieval, observability, and vendor costs can become significant. This role will build visibility into those economics and help leadership understand the cost of AI by application, model, workflow, business unit, or user population.

The successful candidate will combine the mindset of an SRE leader, AI platform operator, technology strategist, incident commander, governance partner, and business operator.

You should be comfortable discussing an LLM evaluation framework with AI engineers, investigating a production incident with platform teams, reviewing token economics with Finance, challenging an AI vendor’s SLA, discussing data controls with Security, and presenting AI operational health to senior leadership.

Ultimately, this role answers one critical enterprise question:

How do we make AI dependable enough to become part of how the company actually operates?


Essential Duties and Responsibilities

AI OPERATING DOMAIN 01

Production AI Reliability

Establish operating standards for production AI applications and services.

  • Define AI production-readiness requirements.
  • Establish service-level objectives and operational expectations.
  • Monitor availability, latency, throughput, errors, and capacity.
  • Develop escalation paths for critical AI services.
  • Improve operational resilience.
  • Coordinate high-availability and disaster-recovery planning where appropriate.
  • Identify recurring reliability problems and drive permanent corrective actions.

Treat AI reliability as an engineering discipline rather than reactive support.


AI OPERATING DOMAIN 02

LLMOps & Generative AI Operations

Develop operational practices supporting enterprise generative AI.

The environment may involve:

  • Large Language Models
  • Foundation-model APIs
  • Enterprise copilots
  • Retrieval-Augmented Generation
  • Vector databases
  • Prompt-management systems
  • AI gateways
  • Model routers
  • Evaluation frameworks
  • Embedding services
  • Fine-tuned models
  • Open-source models

Establish consistent processes for configuration, deployment, evaluation, monitoring, rollback, and lifecycle management.


AI OPERATING DOMAIN 03

Agentic AI Operations

Develop operating controls for increasingly autonomous AI systems.

Support environments involving:

  • AI agents
  • Multi-agent workflows
  • Tool-calling systems
  • Automated research
  • Intelligent workflow orchestration
  • AI-driven business processes

Establish controls around:

Identity → Permissions → Tool Access → Data Access → Action Boundaries → Human Approval → Logging → Evaluation

Monitor agent completion rates, failures, unexpected behaviors, escalations, and downstream business impact.


AI OPERATING DOMAIN 04

AI Observability

Build comprehensive visibility into AI system behavior.

Develop dashboards and monitoring for:

  • Availability
  • Response latency
  • Token consumption
  • Inference performance
  • Error rates
  • Model usage
  • Retrieval quality
  • Evaluation scores
  • Hallucination indicators
  • Safety events
  • User feedback
  • Agent success rates
  • Cost per request
  • Cost per workflow

Integrate AI observability with broader enterprise monitoring wherever practical.


AI OPERATING DOMAIN 05

AI Quality & Evaluation

Partner with AI Engineering, Data Science, Product, and business teams to establish repeatable evaluation processes.

Support:

  • Golden datasets
  • Regression evaluations
  • Human evaluation
  • LLM-as-judge methodologies with appropriate validation
  • Task-completion measurement
  • Groundedness evaluation
  • Retrieval evaluation
  • Safety testing
  • Red-team findings
  • Production feedback loops

Prevent model, prompt, retrieval, or configuration changes from reaching production without appropriate validation.


AI OPERATING DOMAIN 06

Incident Management & AI Command

Establish a mature incident-management model for AI services.

Coordinate response to events involving:

  • AI service outages
  • Model-provider disruptions
  • Significant latency degradation
  • Unexpected model behavior
  • Retrieval failures
  • Data exposure concerns
  • Safety-control failures
  • Excessive cost consumption
  • Agent malfunction
  • Third-party AI service incidents

Lead structured incident response.

Ensure significant incidents produce:

Containment → Root Cause → Corrective Action → Lessons Learned → Preventive Controls


AI OPERATING DOMAIN 07

AI Governance & Responsible Operations

Translate enterprise AI policies into practical operational controls.

Partner with:

  • Cybersecurity
  • Privacy
  • Legal
  • Compliance
  • Risk
  • Data Governance
  • Internal Audit

Support controls involving:

  • Model approval
  • Application registration
  • Data classification
  • AI access
  • Sensitive-data handling
  • Human oversight
  • Logging
  • Third-party AI
  • Model documentation
  • Change control
  • Retention
  • Responsible AI requirements

Maintain sufficient evidence for internal reviews and audits.


AI OPERATING DOMAIN 08

Model & Provider Lifecycle Management

Establish operational discipline across multiple model providers.

Evaluate:

  • Model quality
  • Availability
  • Latency
  • Cost
  • Context capability
  • Security requirements
  • Data policies
  • Geographic availability
  • Rate limits
  • Enterprise support
  • Deprecation schedules

Develop contingency strategies where critical applications depend heavily on a single provider or model.


AI OPERATING DOMAIN 09

AI FinOps & Unit Economics

Create financial visibility into enterprise AI consumption.

Track costs associated with:

  • Tokens
  • Model APIs
  • GPU infrastructure
  • Compute
  • Vector storage
  • Embeddings
  • AI platforms
  • Observability
  • Data processing
  • Third-party AI tools

Develop metrics such as:

Cost per User | Cost per Query | Cost per Workflow | Cost per Agent Task | Cost per Application

Identify optimization opportunities without materially degrading quality.

Partner with Finance and Technology leadership on AI budgeting and forecasting.


AI OPERATING DOMAIN 10

Enterprise AI Platform Operations

Support shared infrastructure used by multiple AI teams.

Potential capabilities may include:

  • Model gateways
  • API management
  • Prompt registries
  • Vector databases
  • Feature stores
  • Model registries
  • Evaluation platforms
  • Secrets management
  • Identity services
  • AI observability
  • Guardrail services
  • Data connectors

Improve platform reliability while enabling engineering teams to deploy AI capabilities faster.


AI OPERATING DOMAIN 11

Security & Access Management

Partner with Security teams to implement appropriate AI controls.

Support:

  • Role-based access
  • Least-privilege principles
  • Service identities
  • Secrets management
  • API-key governance
  • Data-access restrictions
  • Environment separation
  • Audit logging
  • Vendor security review

Ensure employees and AI agents receive appropriate—not unlimited—access to enterprise resources.


AI OPERATING DOMAIN 12

AI Change & Release Management

Establish change processes appropriate for AI systems.

Manage operational risk associated with changes to:

  • Models
  • Prompts
  • System instructions
  • Retrieval pipelines
  • Embeddings
  • Guardrails
  • Tools
  • Data sources
  • Agent permissions
  • Model parameters
  • Provider configurations

Create appropriate testing, approval, deployment, and rollback procedures.


AI OPERATING DOMAIN 13

Enterprise AI Adoption

Partner with business and transformation teams to improve responsible adoption.

Identify:

  • Underutilized AI investments
  • Workflow opportunities
  • User friction
  • Training requirements
  • Repetitive tasks suitable for AI
  • High-value enterprise use cases

Help distinguish meaningful productivity improvements from technology experimentation without measurable value.


AI OPERATING DOMAIN 14

Vendor & Commercial Management

Manage operational relationships with AI technology providers.

  • Review service performance.
  • Track contractual commitments.
  • Escalate production issues.
  • Evaluate enterprise support.
  • Support renewal decisions.
  • Monitor product roadmaps.
  • Identify vendor concentration risk.
  • Partner with Procurement, Legal, Security, and Finance.

Provide leadership with evidence-based recommendations about AI technology investments.


AI OPERATING DOMAIN 15

Team & Operating Model Leadership

Build a disciplined AI Operations function.

Depending on organizational structure, lead or coordinate professionals across:

  • AI Operations
  • MLOps
  • LLMOps
  • Platform Engineering
  • SRE
  • AI Support
  • FinOps
  • AI Governance Operations

Establish clear ownership across Engineering, Product, Data Science, Security, and Operations.

Eliminate ambiguity around who owns production AI after deployment.


Job Qualifications and Requirements

  • 8+ years of progressive experience across technology operations, cloud engineering, platform engineering, SRE, DevOps, MLOps, machine-learning infrastructure, AI platforms, or related disciplines.
  • 3+ years of leadership, technical leadership, or major operational ownership within complex technology environments.
  • Demonstrated understanding of production AI and machine-learning systems.
  • Practical knowledge of generative AI and LLM architecture.
  • Experience supporting cloud-native environments.
  • Strong understanding of production monitoring and observability.
  • Experience with incident management and root-cause analysis.
  • Familiarity with model lifecycle and deployment practices.
  • Understanding of APIs, distributed systems, data pipelines, and modern cloud architectures.
  • Experience implementing operational controls in complex enterprise environments.
  • Strong understanding of security and access-management principles.
  • Demonstrated vendor-management capabilities.
  • Strong financial and operational judgment.
  • Experience communicating technical risk to senior leadership.
  • Bachelor’s degree in Computer Science, Engineering, Information Systems, Data Science, or a comparable technical discipline preferred.

Highly Valuable Technical Exposure

Experience with several of the following is particularly valuable:

  • AWS
  • Microsoft Azure
  • Google Cloud Platform
  • Kubernetes
  • Docker
  • Python
  • Terraform
  • Datadog
  • Grafana
  • Prometheus
  • OpenTelemetry
  • MLflow
  • Databricks
  • Snowflake
  • Vector databases
  • Model gateways
  • LLM observability
  • RAG architectures
  • Model evaluation
  • MLOps platforms
  • CI/CD
  • Infrastructure as Code

Experience operating enterprise AI environments involving leading commercial or open-source model ecosystems is strongly valued.


Personal Capabilities and Qualifications

Operational Judgment

You know that a technically impressive AI system is not necessarily a production-ready system.

Systems Thinking

You understand dependencies across models, prompts, retrieval, data, infrastructure, APIs, security controls, and business workflows.

Incident Leadership

You remain structured and decisive when production behavior becomes uncertain.

AI Curiosity

You continuously study new models, platforms, agent frameworks, operational patterns, and AI capabilities without treating every new technology as automatically production-worthy.

Risk Intelligence

You can distinguish acceptable experimentation from risk requiring immediate intervention.

Financial Awareness

You understand that AI architecture decisions can have significant unit-economic consequences.

Executive Communication

You can explain complex AI operational issues without requiring executives to understand every technical implementation detail.

Cross-Functional Influence

You can align engineers, data scientists, security professionals, business leaders, Finance, Legal, and vendors around shared operational standards.

Continuous Improvement

You prefer eliminating recurring operational problems over repeatedly responding to them.


Strategic Support

The AI Operations Manager will provide critical operational leadership across the enterprise AI portfolio.

Enterprise AI Strategy

Provide operational input before major AI investments move into production.

AI Governance

Translate policies into practical engineering and operational controls.

AI Transformation

Help business functions move from isolated pilots to sustainable production workflows.

Architecture

Provide operational perspectives on model, platform, provider, and infrastructure decisions.

Cybersecurity

Strengthen access, data, identity, logging, and third-party AI controls.

Financial Planning

Provide visibility into AI consumption, capacity, and unit economics.

Vendor Strategy

Help determine where to build, buy, partner, or maintain multiple model-provider options.

Workforce Productivity

Measure whether AI adoption is producing meaningful operational improvement.

Business Continuity

Ensure critical AI-enabled workflows have appropriate resilience and recovery strategies.

The role will help leadership connect:

AI Ambition → Production Reality → Operational Control → Business Value


Working Conditions

  • Hybrid or flexible working arrangements may be available depending on organizational requirements.
  • Regular collaboration with AI Engineering, Machine Learning, Platform Engineering, Cloud, Data, Product, Security, Risk, Legal, Finance, and business teams.
  • Participation in production incident response may occasionally require availability outside standard business hours.
  • Collaboration across multiple geographic regions and time zones may be required.
  • Limited domestic or international travel may be necessary for strategic planning, vendor meetings, leadership sessions, or major technology initiatives.
  • The role involves access to confidential technical architecture, AI configurations, security information, enterprise data strategies, vendor arrangements, and business-sensitive AI initiatives.
  • The operating environment is fast-moving and requires balancing experimentation with enterprise reliability and control.

Job Function

Primary Function: Artificial Intelligence Operations & Platform Management

Operating Scope: Production AI + LLMOps + Reliability + Governance + FinOps

Core Expertise:

AI Operations | Generative AI | LLMOps | MLOps | AI Platforms | AI Agents | RAG | Model Evaluation | AI Observability | SRE | Incident Management | Cloud Infrastructure | AI Governance | Responsible AI | AI FinOps | Vendor Management | Enterprise AI Transformation


Compensation & Benefits

Compensation Package

$220,000 – $255,000 annually

Final compensation will consider AI operations expertise, technical depth, production-scale experience, leadership scope, cloud and platform experience, generative AI knowledge, enterprise complexity, geographic considerations, and overall qualifications.

The broader total rewards package may include:

  • Annual performance incentives
  • Equity or long-term incentives where applicable
  • Comprehensive medical, dental, and vision coverage
  • Employer-sponsored life and disability coverage
  • Retirement savings with employer contributions
  • Generous paid time off
  • Paid parental and family leave
  • Flexible or hybrid working arrangements
  • Professional AI and cloud certifications
  • Technical conference participation
  • Advanced learning and leadership-development resources
  • Enterprise AI experimentation resources
  • Wellness and employee-assistance programs
  • Additional benefits according to organizational policy

Why Join Us

The next stage of enterprise AI will not be defined only by who has access to the most capable model.

It will be defined by who can make AI reliable enough, secure enough, measurable enough, economical enough, and trusted enough to operate at scale.

That is the problem this role is built to solve.

You will work beyond the experimental layer of AI.

You will help determine what happens after the prototype works.

How is it monitored?

How do we know the answers remain good?

What happens when the model changes?

What happens when an agent receives access to enterprise systems?

How quickly can we detect abnormal behavior?

What does each AI workflow actually cost?

How do we recover when something fails?

And how do we scale from a successful pilot to thousands of users without losing control?

Those questions will become increasingly important as AI moves deeper into everyday business operations.

You will have the opportunity to shape the operating model before many of those patterns become permanently established.

The mission is to move enterprise AI from:

Experiment → Deployment → Reliability → Governance → Scale → Measurable Value

And ultimately build the connection between:

Models + Data + Infrastructure + Controls + People → Trusted AI Operations

For a technology leader who wants to work where AI innovation meets production reality, this role offers the opportunity to build the operational foundation that allows an enterprise to use artificial intelligence with confidence.