Deploy a dedicated L40 cloud GPU on EmpirioLabs AI. One-click JupyterLab, ComfyUI, vLLM serving, and web terminal templates, billed per second only while the instance runs.
Every instance launches from a one-click template, with no image building or driver setup:
Per second at the listed hourly rate, only while the instance is running. The rate is locked in when you deploy, and stopping or destroying the instance stops the charge.
One-click templates cover JupyterLab notebooks, ComfyUI, vLLM model serving with any Hugging Face model id, and a browser web terminal for your own workloads.
The facts above show live availability from the deploy catalog. When this GPU is out of stock, other GPUs with similar VRAM are usually available on the GPU Cloud page.
Yes. A one-click model instance includes a chat page right on the instance view, and the served endpoint is OpenAI-compatible for use from your code. Ollama instances expose the standard Ollama API.
Yes. Deploy, stop, and destroy instances under /v1/gpu on api.empiriolabs.ai, and reach the running workload through the authenticated EmpirioLabs connect endpoint. The full reference is in the GPU Cloud docs.
Check out our pricing or reach out if you want your own model deployed on our stack.