Beta · Linux + NVIDIA · Apple Silicon

Your GPUs. One private AI cloud.

You already have the hardware. Nodeau turns the GPUs and computers you own into one private AI cloud. Run models on them, add machines as you go, and call everything through the OpenAI-compatible API your apps already speak.

Free on one machine. No account needed to start.

An example Nodeau fleet: three machines with four NVIDIA GPUs between them. One serves a chat model, one serves an embedding model, and one splits a large model across its two cards. Each model has an OpenAI-compatible endpoint.

Why Nodeau

The hard part was never the GPU.

The card in your desktop can run a genuinely useful model today. Getting there usually means drivers, containers, model formats and a lot of memory math. Nodeau does that work for you.

It checks each model fits before it starts, downloads and verifies the weights, and hands you an endpoint that works. Start with the machine in front of you. Add another whenever you like, and they all become one fleet you can see from anywhere.

Runs on hardware you ownYour GPUs do the work, and your prompts stay with them.
Checks the fit firstEvery model lands on a GPU that can hold it.
OpenAI-compatiblePoint the client you already use at it.
Private by defaultBound to 127.0.0.1, with a key made on your machine.

Getting started

Three commands to your first model.

Most of the time goes to downloading the model. On a Mac it's even simpler: no drivers and no password.

Install Nodeau

One small program, checksum verified, into a folder you own.

curl -fsSL https://get.nodeau.ai/install.sh | bash

Set up the machine

Nodeau looks the machine over, shows you its plan and every command that needs sudo, and asks before it changes anything.

nodeau install

Start a model

It offers a starter model, checks that it fits your GPU, downloads and verifies it, then starts it.

nodeau quickstart

Use it

Point any OpenAI-compatible client at http://127.0.0.1:8080/v1. That's it. Your model is live!

On a Mac, step three is nodeau run qwen3.5-4b-q4km. The full guide

nodeau quickstart
Recommended starter model

  Qwen3.5-4B Q4_K_M

  Download
    ~2.55 GiB

  License
    Apache-2.0

  Downloaded from
    huggingface.co

  Nodeau checked your graphics card and expects this to fit.

==> downloading Qwen3.5-4B-Q4_K_M.gguf
==> waiting for Nodeau to admit and start it
==> starting the local endpoint

──────────────────────────────────────
Nodeau is ready.
──────────────────────────────────────

Model
  qwen3.5-4b-q4km

Local API
  http://127.0.0.1:8080/v1

  Only this computer can reach it.

Your hardware

Bring the hardware you already have.

NVIDIA cards, Apple Silicon Macs, big cards and small ones. Nodeau learns what each GPU can really give and places every model on one that can actually run it.

NVIDIA GPUs on Linux

Ubuntu 24.04 with a working NVIDIA driver. Nodeau has run on the RTX 5070 Ti, 5060 Ti, 3080 and 2080, and it estimates carefully for cards it's meeting for the first time.

Apple Silicon Macs

Chat, embeddings, reranking and vision run natively on the Mac's own GPU through Metal, several models at once. No drivers, no password. A Mac runs standalone as your own private AI endpoint.

Several GPUs in one machine

Every card is scheduled on its own. Run a different model on each, or split a model that's too big for one card across two.

Different cards, one fleet

Mix sizes and generations. Each machine reports what it really has, and Nodeau matches every model to a GPU that fits it.

A measurement when Nodeau has one, a careful estimate otherwise, and it always tells you which. When a model is too big for a card, you get the numbers and a suggestion instead of a crash.

What you can run

Much more than a chat box.

Pick a model from the catalog, or bring your own. The same engine serves chat, search, ranking, structured data and images, and the same fit check guards all of it.

Chat and completion

Streaming included, and it works from any OpenAI client.

Embeddings

Turn text into vectors for search, RAG and clustering, on your own GPU.

Reranking

Score documents against a query so the best answer comes first.

Structured output

Ask for JSON that follows your schema, and get exactly that shape back.

Tool calling

Let a model decide when to call your functions, with arguments your code can parse.

Vision

Drop in an image and ask a model what it sees.

Batch inference

Hand over a file of requests. Your GPUs work through it and every result comes back matched to its input.

Your own models

Have a GGUF you love? Import it, and Nodeau tests what it can really do on your card before you rely on it.

Chat, embeddings, reranking and vision run everywhere Nodeau runs. Structured output, tool calling, batch jobs and your own models run on Linux machines with NVIDIA GPUs.

8 GB and up
  • Qwen3.5-4B starter
  • Gemma 4 E4B vision
  • Qwen3 Embedding 0.6B
  • BGE Reranker v2 m3
12 GB and up
  • Qwen3.5-9B
  • Gemma 4 12B vision
16 GB and up
  • GPT-OSS 20B mixture of experts
24 GB or two cards
  • Qwen3.8-27B flagship

Browse the catalog

One fleet

Add another machine when you're ready.

Run nodeau fleet invite on the machine you have and nodeau join on the new one. That's the whole thing. The new machine shows up in your fleet and Nodeau starts placing work on it.

  • See everything in one place. Every machine, card and model, from any browser.
  • Know how each machine is doing. Processor, memory, storage, network and GPU, with alerts when something needs you.
  • Set your own guardrails. Which models, which cards, which machines, and how much can run at once.
  • See what your hardware did. Usage in accelerator-hours, and a record of who changed what.
  • Update the fleet from one place. Preview the rollout first, then Nodeau moves through your machines one at a time and checks each one before the next.
your machines→ connect out →Nodeau Cloud

Your machines start every conversation with Nodeau Cloud, which keeps your network boundary simple. Models keep serving whether or not the cloud is there.

nodeau fleet list
MY NODEAU
  fleet     2 of 3 machines (home-pro)

  garage   online   worker
    NVIDIA GeForce RTX 3080  9,877 MiB

  studio   online   control-plane
    NVIDIA GeForce RTX 5060 Ti  15,827 MiB

  Add a machine:  nodeau fleet invite
nodeau health
garage  (worker, ready, reported 6s ago)
  CPU           0 %    16 threads
  memory        12 %   1.9 GiB of 15.5 GiB in use
  storage used  21 %   709.7 GiB free
  network       7.5 KiB/s in, 2.9 KiB/s out
  GPU           RTX 3080, 44 of 9,877 MiB used, healthy
python
import os
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:8080/v1",
    api_key=os.environ["NODEAU_API_KEY"],
)

reply = client.chat.completions.create(
    model="qwen3.5-4b-q4km",
    messages=[
        {"role": "user", "content": "Hello from my GPU!"}
    ],
)
print(reply.choices[0].message.content)

For developers

Change one line and your app runs on your hardware.

Nodeau speaks the OpenAI API, so the SDKs and tools you already use just work. Set the base URL, use the key Nodeau made for you, and go.

  • Chat, streaming, embeddings and reranking over plain HTTP
  • Every command that prints a table also speaks --json
  • Exit codes that tell a refusal apart from a model that's still loading
  • nodeau placement explain shows why a model landed where it did
  • A local dashboard, served by Nodeau on your own machine

Read the API guide

Plans

Start free. Grow when you're ready.

Your hardware does the work, so nothing is metered by the token. Plans are about how many machines you run and how you run them.

Home

One machine, one GPU.

Free

A great way to turn a desktop or a Mac into your own AI endpoint.

Get Nodeau

Business

Your organisation's private AI fleet.

Let's talk

As many machines and GPUs as you run, with business support and a hand planning the rollout.

Talk to us

Compare the plans

Your GPUs are already capable. Put them to work!

About ten minutes from here to a model answering on your own hardware. Got a room full of machines instead? We'd love to hear about it.