Rent H100 NVL by the hour

Deploy a dedicated H100 NVL cloud GPU on EmpirioLabs AI. One-click JupyterLab, ComfyUI, vLLM serving, and web terminal templates, billed per second only while the instance runs.

Specs and price

VRAM
94 GB
Price
$4.20/hr
Availability (Live)
Unavailable
Multi-GPU configs
1x
Deploy H100 NVL

What you can run on it

Every instance launches from a one-click template, with no image building or driver setup:

Serve an LLM (vLLM)
Deploy a ready-made model card (gpt-oss, Qwen3.5 and 3.6, Qwen3 Coder) or paste any Hugging Face model id. The launch is auto-configured for the model and this GPU, and it serves an OpenAI-compatible endpoint with a chat page built right into the instance view.
Run models with Ollama
Pull any open model, including GGUF quantizations, and use it through the Ollama API. Good for quantized weights and the Ollama tooling you already know.
JupyterLab
A CUDA PyTorch notebook environment, open in the browser in minutes.
ComfyUI
The node-based image and video workflow UI, with the Manager panel for installing custom nodes and models.
Web terminal
A browser bash shell into the instance for your own Docker-based workloads and serving stacks.

How GPU Cloud works

1. Pick a configuration
Choose the GPU count and a runtime storage target from the live catalog.
2. Pick a template
JupyterLab, ComfyUI, a vLLM model server, or a browser web terminal, ready in minutes.
3. Connect
Open the workload in the browser or call it through the authenticated EmpirioLabs connect endpoint.

Common questions

How is it billed?

Per second at the listed hourly rate, only while the instance is running. The rate is locked in when you deploy, and stopping or destroying the instance stops the charge.

What can I run on it?

One-click templates cover JupyterLab notebooks, ComfyUI, vLLM model serving with any Hugging Face model id, and a browser web terminal for your own workloads.

Is it available right now?

The facts above show live availability from the deploy catalog. When this GPU is out of stock, other GPUs with similar VRAM are usually available on the GPU Cloud page.

Can I chat with the model I deploy?

Yes. A one-click model instance includes a chat page right on the instance view, and the served endpoint is OpenAI-compatible for use from your code. Ollama instances expose the standard Ollama API.

Can I manage instances through the API?

Yes. Deploy, stop, and destroy instances under /v1/gpu on api.empiriolabs.ai, and reach the running workload through the authenticated EmpirioLabs connect endpoint. The full reference is in the GPU Cloud docs.

Ready to use better endpoints?

Check out our pricing or reach out if you want your own model deployed on our stack.