Deployment architecture guide

Choose an Azure endpoint around the inference workload

Azure can host inference workloads with different requirements for responsiveness, scale and implementation control. A hosted model interface and a custom deployment have distinct contracts: neither should be treated as a universal tensor API. Start by deciding whether callers need an immediate response, whether the model consumes raw data or numerical features and who owns preprocessing. Then select the deployment approach that supports those decisions, keeping the public application contract separate from the resources that execute the model.

Match the endpoint to the caller

An online inference workflow returns a result within the request path. A batch workflow starts asynchronous work and needs a way to track completion and collect results. These execution models create different obligations for timeouts, retries and partial failure. Describe those obligations before comparing deployment settings.

For a custom model, identify the scoring code, model artifacts and runtime environment that must travel together. For a hosted language interface, inspect the supported request and response contract rather than assuming access to internal tensors. Separate the stable endpoint interface from the selected deployment. This makes it easier to evaluate a replacement implementation without forcing every caller to understand hardware, model packaging or changes in the internal serving stack.

Document the data before scaling it

A numerical request needs more than an array field. Specify element type, axis meaning, accepted shapes, feature order and the treatment of missing values. For text, define encoding, size limits and the preprocessing configuration. Keep record identifiers attached to outputs, especially when work is split into batches.

  • Version the model, tokenizer or feature extractor, and label mapping together.
  • Reject inconsistent shapes before expensive model allocation or execution.
  • Record normalization rules so training and serving use the same interpretation.
  • Measure serialization, transfer, preprocessing and inference separately.

Choose representative small and large requests for verification. A payload that passes a JSON parser may still violate the model’s semantic contract, while a faster accelerator cannot fix time spent decoding oversized inputs.

Promote a complete, observable release

Compare a candidate deployment with reviewed examples and expected workload patterns. Measure prediction quality and operational behavior separately. A response can arrive quickly with a valid shape and still contain the wrong category mapping. Keep the model and preprocessing versions in execution metadata so unexpected changes can be traced.

Define promotion and rollback criteria in advance. Retain the preceding model, environment and transformation artifacts until the replacement passes those checks. Where the selected endpoint supports it, a limited traffic allocation or an isolated comparison path can expose realistic behavior. For batch work, verify that failed partitions can be identified and retried without losing record correspondence. Finish by testing the path with the application’s actual identity and required data dependencies, not only a developer’s account.

Official referenceMicrosoft Learn: Online endpoints for real-time inference

Use the official reference for the documented interface; the workflow recommendations above are Tensor API Lab guidance.