Self-hosted model operations

Container and On-Premises Model Service Runbook

Operate a shared containerized model service with pinned artifacts, explicit host and cluster boundaries, measured capacity, and reversible releases.

Core concept
CurrentNext review due 2026-09-25Content version 0.42.0
Format
Playbook
Level
Advanced
Audience
Developer, Operator, Leader
Owner
Project42 Editorial
Review cadence
Every 60 days
Prerequisites
An approved service contract, model artifact, image, and dependency manifest; Named host, cluster, network, identity, storage, and incident owners; Representative evaluation and load profiles with service and recovery objectives
01

Construct one verifiable serving unit

Treat the release as a complete unit: model and tokenizer digests, license and provenance decision, immutable image, runtime and driver compatibility, templates and adapters, gateway contract, identity and authorization policy, network and egress rules, secret references, resource requests and limits, storage, telemetry, evaluation set, load profile, rollout and rollback policy, and runbook. A model tag or container image alone cannot reproduce service behavior.

Separate serving and management paths. Grant deployment, secret, cluster, and model-administration authority only to named roles. Run the workload with the minimum operating-system and container privileges, constrain host mounts and device access, protect the runtime or orchestrator control plane, and keep secrets and learner content out of images, manifests, logs, and screenshots.

02

Operate within measured capacity and failure domains

Declare context, output, body, concurrency, rate, queue, deadline, retry, accelerator, memory, storage, and cost boundaries. Exercise load through the real gateway and serving path, including warm-up, long requests, cancellation, overload, dependency loss, instance loss, and telemetry loss. Resource requests and limits are scheduling and isolation inputs; they are not proof of model throughput or safe queue behavior.

Use bounded admission and backpressure before saturation. Scale only on signals that explain the bottleneck, and verify the full replica lifecycle: artifact identity, readiness, traffic eligibility, drain, termination, and replacement. Keep at least one compatible known-good release manifest and prove that its artifacts and dependencies remain available.

Container and on-premises release record
text
Task: [SHARED MODEL-SERVICE WORKLOAD]
Scope: [HOSTS/CLUSTER, USERS, DATA, NETWORKS, STORAGE, TOOLS, AND SERVICE OBJECTIVES]
Permissions: [MODEL/LICENSE, IMAGE REGISTRY, DEPLOY, CLUSTER, SECRET, NETWORK, AND RELEASE AUTHORITY]
Exact build: [MODEL/TOKENIZER DIGESTS, IMAGE DIGEST, RUNTIME/DRIVER, ADAPTER, POLICY, CONFIG, AND EVAL VERSION]
Isolation: [USER, CAPABILITIES, MOUNTS, DEVICES, NAMESPACES, INGRESS, EGRESS, AND MANAGEMENT PATH]
Capacity: [REQUEST/LIMIT, CONTEXT, OUTPUT, CONCURRENCY, RATE, QUEUE, DEADLINE, AND FAILURE DOMAIN]
Verification: [PROVENANCE, IDENTITY, NEGATIVE ACCESS, REPRESENTATIVE EVAL, LOAD, INSTANCE LOSS, DRAIN, AND RESTORE]
Stop conditions: [ARTIFACT/POLICY DRIFT, CONTROL-PLANE EXPOSURE, CRITICAL EVAL FAILURE, UNBOUNDED QUEUE, SATURATION, OR FAILED ROLLBACK]
Recovery: [STOP ROUTING, DRAIN/CANCEL, RECONCILE EFFECTS, RESTORE COMPLETE KNOWN-GOOD UNIT, VERIFY POSTCONDITIONS]
03

Expected evidence and verification

Expected evidence includes the complete release manifest, signatures or digests, provenance and license disposition, runtime compatibility, least-privilege and network tests, secret-source evidence, representative evaluation and load results, queue and overload behavior, replica and failure-domain behavior, telemetry and alerts, backup or artifact-retention proof, rollback rehearsal, owners, and review date.

Stop new traffic when artifacts or policy drift, a management interface becomes exposed, a critical evaluation or security case fails, queues lose their bound, the service saturates outside its decision, or the known-good release cannot be restored. Preserve secret-free evidence, contain the narrow boundary, reconcile in-flight and external effects, restore the full compatible unit, and verify identity, authorization, quality, capacity, telemetry, and user-visible postconditions.