Self-hosted model operations
Self-Hosted Model Deployment Shape Decision
Choose among workstation, container or on-premises, edge, and cloud-managed model compute using workload, authority, evidence, and recovery requirements.
- Format
- Decision Path
- Level
- Advanced
- Audience
- Developer, Operator, Leader
- Owner
- Project42 Editorial
- Review cadence
- Every 60 days
- Prerequisites
- A bounded workload and prohibited-use statement; Data classification, service objectives, and recovery objectives; Named owners for model, infrastructure, security, and release decisions
Compare responsibility, not product labels
A workstation can minimize network dependence and support individual experimentation, but it inherits device access, thermal, power, storage, patching, and single-machine recovery limits. A container or on-premises service can standardize runtime and shared access, while adding host, image, network, scheduler, capacity, backup, and incident ownership. An edge service can reduce round-trip latency and operate with intermittent connectivity, but must fit device memory, power, update, observability, physical-access, and rollback constraints. Cloud-managed compute can delegate hardware provisioning and some service operations while retaining model selection, authorization, data, evaluation, quota, cost, lifecycle, and recovery responsibilities.
Do not select a shape from one benchmark or a privacy slogan. Compare the same exact model, runtime contract, request distribution, data boundary, failure cases, and service objectives. A shape is viable only when the organization has authority to use the artifacts and environment, can verify the intended workload, can stop unsafe traffic, and can restore a known-good serving unit.
Record an expiring, evidence-backed choice
Evaluate workload locality, disconnected operation, latency distribution, concurrency, context and output bounds, accelerator availability, energy, idle capacity, portability, isolation, staffing, support, update cadence, telemetry, cost, and recovery. Date hardware, pricing, quota, version, and lifecycle observations or link the decision to a current first-party source instead of copying volatile values into a long-lived guide.
Keep the stable service contract separate from each environment adapter. The contract identifies inputs, outputs, errors, limits, identity, evidence, and recovery; adapters hold workstation, container, edge, or cloud-specific implementation details. This makes a future move a measured adapter change instead of a rewrite of the operating promise.
Task: [BOUNDED MODEL-SERVICE WORKLOAD]
Scope: [USERS, DATA, LOCATIONS, REQUEST SHAPE, AND NON-GOALS]
Permissions: [MODEL/LICENSE, DATA, DEVICE/CLOUD, NETWORK, AND RELEASE AUTHORITY]
Exact build: [MODEL REVISION + DIGEST, RUNTIME/IMAGE, ADAPTER, POLICY, AND CONFIG]
Candidates: [WORKSTATION | CONTAINER/ON-PREMISES | EDGE | CLOUD-MANAGED]
Demand: [CONTEXT, OUTPUT, CONCURRENCY, BURST, LATENCY, AVAILABILITY, AND COST]
Control boundary: [IDENTITY, NETWORK, SECRETS, TELEMETRY, UPDATE, AND SUPPORT]
Decision evidence: [QUALITY/SAFETY, LOAD, FAILURE, TCO, STAFFING, AND PORTABILITY]
Decision owner and review date: [AUTHORIZED ROLE + YYYY-MM-DD]
Verification: [REPRODUCE BUILD, REPRESENTATIVE TESTS, FAILURE TEST, AND RESTORE]
Stop conditions: [LICENSE GAP, DATA VIOLATION, CRITICAL EVAL FAILURE, UNSAFE SATURATION, OR UNOWNED RECOVERY]
Recovery: [STOP ADMISSION, RECONCILE WORK, RESTORE KNOWN-GOOD UNIT, VERIFY POSTCONDITIONS]Expected evidence and verification
Expected evidence includes the bounded mission, data and authority boundaries, exact artifact and runtime identities, one comparable evaluation and load profile, failure and restore results, cost assumptions, operating owners, residual risks, selected shape, rejected alternatives, and next review date. A successful chat response is not deployment evidence.
Verify the selected shape under representative load and at least one declared failure. Stop when a legal, data, identity, critical-evaluation, capacity, or recovery gate fails. Preserve secret-free evidence, block new work, reconcile in-flight effects, restore the verified fallback, and reopen the decision when demand, hardware, pricing, quota, runtime support, or organizational ownership materially changes.