24/06/2026
Choosing between AI performance and strict data compliance? Thatโs all water under the bridge now. ๐ Exoscale Dedicated Inference is officially out of preview and live for production-grade workloads. ๐
You can now deploy any Hugging Face model as a production-ready, OpenAI-compatible API endpoint, backed by the security of a fully sovereign European infrastructure AND powered by dedicated NVIDIA GPUs
What this means for your production environment:
๐๐ฏ๐๐ผ๐น๐๐๐ฒ ๐๐๐ผ๐น๐ฎ๐๐ถ๐ผ๐ป:
Dedicated instances mean zero resource sharing. Your proprietary data and prompts never leave your environment.
๐๐ฟ๐ถ๐ฐ๐๐ถ๐ผ๐ป๐น๐ฒ๐๐ ๐๐ป๐๐ฒ๐ด๐ฟ๐ฎ๐๐ถ๐ผ๐ป:
Swap in your preferred models and start querying immediately through standard API frameworks.
๐๐ป๐๐ฒ๐ฟ๐ฝ๐ฟ๐ถ๐๐ฒ ๐ฆ๐๐ฎ๐ฏ๐ถ๐น๐ถ๐๐:
Production workloads are now fully covered by a 99.95% SLA, with operational and support processes fully integrated into the standard Exoscale service lifecycle.
๐ฉ๐ฒ๐ฟ๐๐ฎ๐๐ถ๐น๐ฒ ๐๐ ๐๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ:
Fully optimized not just for LLMs, but also for complex embeddings, RAG applications, AI agents and custom inference APIs.
Our documentation has been completely expanded with deployment scaling guidance, updated CLI references service boundaries and model compatibility requirements.
Learn more: https://changelog.exoscale.com/en/ai-dedicated-inference-is-now-generally-available