# LLM-Gate

A local language model with a door on it. Two containers and a folder of models, depending on nothing else:
copy this directory to another machine, run `./gate up`, and the same models answer on the same port with the
same keys.

```
LLM-Gate/
├── docker-compose.yml      the two containers
├── gate  (gate.cmd)        the one command (bash; gate.cmd runs it from cmd/PowerShell)
├── .env                    ports, resources, where the models live   (.env.example beside it)
├── keys.conf               who may call, one line per caller          (keys.conf.example beside it)
├── templates/              the door's rules (nginx), rendered at start-up
└── models/                 the downloaded models — this is why the directory is worth copying
```

`.env` and `keys.conf` hold settings and secrets; neither is meant to be committed.

## Start

```bash
./gate up                       # creates .env and keys.conf on the first run
./gate keys add my-app          # mint a key for a caller, printed once
./gate pull qwen3:30b-a3b       # download a model into ./models
./gate test                     # one real request, through the door, with a real key
```

`./gate status` shows the containers, the health, the installed models and who has a key.

## Calling it

```bash
curl http://<host>:11444/v1/chat/completions \
  -H "Authorization: Bearer $KEY" -H 'Content-Type: application/json' \
  -d '{"model":"qwen3:30b-a3b","messages":[{"role":"user","content":"hello"}]}'
```

It speaks **Ollama's own API** (`/api/generate`, `/api/chat`, `/api/tags`) and its **OpenAI-compatible** one
(`/v1/chat/completions`, `/v1/models`). An OpenAI-shaped client needs only its base URL set to
`http://<host>:11444/v1/` and its api key set to one of the keys here. `X-API-Key: <key>` works too, for a curl
written by hand.

| Path | Key? | |
|---|---|---|
| `/healthz` | no | `ok`, or `503 unavailable` when the model server cannot be reached. It names nothing — not the product, not its version, not the models — so a port scan learns nothing from it. It is also what the container's own healthcheck reads. |
| everything else | **yes** | `401` without a key that is in `keys.conf` |
| `/api/pull`, `/api/push`, `/api/create`, `/api/delete`, `/api/copy`, `/api/blobs` | — | `403` **for everybody, key or no key** |

That last row is the reason this project exists in front of Ollama rather than beside it. Ollama has no
authentication of its own at any setting, and its API can download, delete and create models as well as answer
questions. Published as it comes, on a network, that is handed to whoever is on that network. Changing what
models exist is `./gate pull` and `./gate rm` on this host, not a caller's business.

## Keys

One per caller, never one shared by several:

```bash
./gate keys                     # who has one (shown as first…last, never in full)
./gate keys add thinkera        # mint, reload the door, print it once
./gate keys rm thinkera         # take it away; the door reloads immediately
./gate logs gate                # <ip> thinkera "POST /v1/chat/completions HTTP/1.1" 200 1843 12.6s
```

The name is what the log carries in place of the key, so «who called, how often, how long did it take» is a
question the log can answer — and a caller can be removed without disturbing the others.

## The models folder

`./models` is a plain directory, not a Docker volume, so the models travel with this one. That is the whole
point of the layout, and it has two consequences worth knowing:

* **Keep this directory off a synced folder** (OneDrive, Dropbox, iCloud). `qwen3:30b-a3b` is about 19 GB and
  `llama4:16x17b` about 67 GB; a sync client will try to upload every byte of it, repeatedly.
* To put the models on another disk, set `MODELS_DIR` in `.env` — but then copying this directory no longer
  carries them, which is the one thing this layout is for.

| Model | Download | Memory while it answers |
|---|---|---|
| `qwen3:0.6b` | ≈ 0.5 GB | ≈ 1 GB — small enough to prove the wiring works |
| `gpt-oss:20b` | ≈ 14 GB | ≈ 18 GB |
| `qwen3:30b-a3b` | ≈ 19 GB | ≈ 25 GB — only a few billion of its parameters run per token, which is what makes it bearable on a processor |
| `llama4:16x17b` | ≈ 67 GB | ≈ 80 GB |

On a laptop `.env`'s defaults are right. On a 48-thread, 251 GB server: `LLM_CPUS=40`, `LLM_MEMORY=160g`,
`LLM_MAX_LOADED_MODELS=2`, `LLM_NUM_PARALLEL=2`.

## Where it is published

`GATE_BIND=0.0.0.0` in `.env` publishes the door on every interface, which is what makes it reachable from
another machine — and why the keys matter. On a host whose every interface faces the internet, consider
`GATE_BIND=127.0.0.1` plus an SSH tunnel, or a firewall rule that only lets the callers you know through.

## Pointing the Thinkera platform at it

```bash
cd <Thinkera-Enterprise>
./tk moodle local/tkai/cli/set_gate.php --url=http://host.docker.internal:11444/v1/ --key=<key>
```

or, as an administrator: *Site administration → Plugins → Local plugins → Thinkera AI*, a connection of type
**«متوافق مع OpenAI»** whose address is the gate's `/v1/` and whose key is one minted here.
`host.docker.internal` is how a container on the same machine reaches this host; from another machine use its
address. The platform's own `curl` is told to ignore Moodle's host and port blocking for these calls, so a
private address and port 11444 are reachable where they would normally be refused.
