Tag: model-serving
Concepts
- Hot-Reloading Models Without Downtime
- In-Process vs RPC Model Inference
- Inference Batching and Throughput
- Packaging and Serialising Model Artefacts
- Model Rollback and Freeze Procedures
- ONNX and Cross-Language Model Export
- Pruning and Model Compression
- Quantisation and Pruning for Latency
- Serving Models Inside a Trading System