Your GPUs. One private AI cloud.
You already have the hardware. Nodeau turns the GPUs and computers you own into one private AI cloud. Run models on them, add machines as you go, and call everything through the OpenAI-compatible API your apps already speak.
Free on one machine. No account needed to start.
Why Nodeau
The hard part was never the GPU.
The card in your desktop can run a genuinely useful model today. Getting there usually means drivers, containers, model formats and a lot of memory math. Nodeau does that work for you.
It checks each model fits before it starts, downloads and verifies the weights, and hands you an endpoint that works. Start with the machine in front of you. Add another whenever you like, and they all become one fleet you can see from anywhere.
Getting started
Three commands to your first model.
Most of the time goes to downloading the model. On a Mac it's even simpler: no drivers and no password.
Install Nodeau
One small program, checksum verified, into a folder you own.
curl -fsSL https://get.nodeau.ai/install.sh | bash
Set up the machine
Nodeau looks the machine over, shows you its plan and every command that needs sudo, and asks before it changes anything.
nodeau install
Start a model
It offers a starter model, checks that it fits your GPU, downloads and verifies it, then starts it.
nodeau quickstart
Use it
Point any OpenAI-compatible client at http://127.0.0.1:8080/v1. That's it. Your model is live!
On a Mac, step three is nodeau run qwen3.5-4b-q4km.
The full guide
Recommended starter model
Qwen3.5-4B Q4_K_M
Download
~2.55 GiB
License
Apache-2.0
Downloaded from
huggingface.co
Nodeau checked your graphics card and expects this to fit.
==> downloading Qwen3.5-4B-Q4_K_M.gguf
==> waiting for Nodeau to admit and start it
==> starting the local endpoint
──────────────────────────────────────
Nodeau is ready.
──────────────────────────────────────
Model
qwen3.5-4b-q4km
Local API
http://127.0.0.1:8080/v1
Only this computer can reach it.
Your hardware
Bring the hardware you already have.
NVIDIA cards, Apple Silicon Macs, big cards and small ones. Nodeau learns what each GPU can really give and places every model on one that can actually run it.
NVIDIA GPUs on Linux
Ubuntu 24.04 with a working NVIDIA driver. Nodeau has run on the RTX 5070 Ti, 5060 Ti, 3080 and 2080, and it estimates carefully for cards it's meeting for the first time.
Apple Silicon Macs
Chat, embeddings, reranking and vision run natively on the Mac's own GPU through Metal, several models at once. No drivers, no password. A Mac runs standalone as your own private AI endpoint.
Several GPUs in one machine
Every card is scheduled on its own. Run a different model on each, or split a model that's too big for one card across two.
Different cards, one fleet
Mix sizes and generations. Each machine reports what it really has, and Nodeau matches every model to a GPU that fits it.
Qwen3.5-9B checked against two cards. On an RTX 5060 Ti it needs 5,800 MiB, a measured figure, out of 15,315 MiB available. On an RTX 3080 it needs 7,210 MiB, an estimate, out of 9,365 MiB available. It fits on both.
A measurement when Nodeau has one, a careful estimate otherwise, and it always tells you which. When a model is too big for a card, you get the numbers and a suggestion instead of a crash.
What you can run
Much more than a chat box.
Pick a model from the catalog, or bring your own. The same engine serves chat, search, ranking, structured data and images, and the same fit check guards all of it.
Chat and completion
Streaming included, and it works from any OpenAI client.
Embeddings
Turn text into vectors for search, RAG and clustering, on your own GPU.
Reranking
Score documents against a query so the best answer comes first.
Structured output
Ask for JSON that follows your schema, and get exactly that shape back.
Tool calling
Let a model decide when to call your functions, with arguments your code can parse.
Vision
Drop in an image and ask a model what it sees.
Batch inference
Hand over a file of requests. Your GPUs work through it and every result comes back matched to its input.
Your own models
Have a GGUF you love? Import it, and Nodeau tests what it can really do on your card before you rely on it.
Chat, embeddings, reranking and vision run everywhere Nodeau runs. Structured output, tool calling, batch jobs and your own models run on Linux machines with NVIDIA GPUs.
- Qwen3.5-4B starter
- Gemma 4 E4B vision
- Qwen3 Embedding 0.6B
- BGE Reranker v2 m3
- Qwen3.5-9B
- Gemma 4 12B vision
- GPT-OSS 20B mixture of experts
- Qwen3.8-27B flagship
One fleet
Add another machine when you're ready.
Run nodeau fleet invite on the machine you have and
nodeau join on the new one. That's the whole thing.
The new machine shows up in your fleet and Nodeau starts placing
work on it.
- See everything in one place. Every machine, card and model, from any browser.
- Know how each machine is doing. Processor, memory, storage, network and GPU, with alerts when something needs you.
- Set your own guardrails. Which models, which cards, which machines, and how much can run at once.
- See what your hardware did. Usage in accelerator-hours, and a record of who changed what.
- Update the fleet from one place. Preview the rollout first, then Nodeau moves through your machines one at a time and checks each one before the next.
Your machines start every conversation with Nodeau Cloud, which keeps your network boundary simple. Models keep serving whether or not the cloud is there.
MY NODEAU
fleet 2 of 3 machines (home-pro)
garage online worker
NVIDIA GeForce RTX 3080 9,877 MiB
studio online control-plane
NVIDIA GeForce RTX 5060 Ti 15,827 MiB
Add a machine: nodeau fleet invite
garage (worker, ready, reported 6s ago)
CPU 0 % 16 threads
memory 12 % 1.9 GiB of 15.5 GiB in use
storage used 21 % 709.7 GiB free
network 7.5 KiB/s in, 2.9 KiB/s out
GPU RTX 3080, 44 of 9,877 MiB used, healthy
import os
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:8080/v1",
api_key=os.environ["NODEAU_API_KEY"],
)
reply = client.chat.completions.create(
model="qwen3.5-4b-q4km",
messages=[
{"role": "user", "content": "Hello from my GPU!"}
],
)
print(reply.choices[0].message.content)
For developers
Change one line and your app runs on your hardware.
Nodeau speaks the OpenAI API, so the SDKs and tools you already use just work. Set the base URL, use the key Nodeau made for you, and go.
- Chat, streaming, embeddings and reranking over plain HTTP
- Every command that prints a table also speaks
--json - Exit codes that tell a refusal apart from a model that's still loading
nodeau placement explainshows why a model landed where it did- A local dashboard, served by Nodeau on your own machine
Plans
Start free. Grow when you're ready.
Your hardware does the work, so nothing is metered by the token. Plans are about how many machines you run and how you run them.
Home
One machine, one GPU.
A great way to turn a desktop or a Mac into your own AI endpoint.
Home Pro
Up to three machines, two GPUs in each.
For home labs and developers with more than one machine. Adds batch inference and remote control of your fleet.
Business
Your organisation's private AI fleet.
As many machines and GPUs as you run, with business support and a hand planning the rollout.
Your GPUs are already capable. Put them to work!
About ten minutes from here to a model answering on your own hardware. Got a room full of machines instead? We'd love to hear about it.