A self-hosted, authenticated inference API serving open-weights models from a private GPU cluster — no data leaves the building, no public API ever sees your prompts. This page is served from the same host as the API. The API itself is key-gated, exactly as it would be in production.
/v1/chat/completions · /v1/modelsThe prototype demonstrates the self-hosted serving pattern — open weights behind an authenticated API, data on-prem. The model behind the router is a configuration change, not an architecture change.
The production recommendation is a non-Chinese frontier model (Mistral Large 4 class) for the sovereignty and licensing posture the entry argues for — the serving stack doesn't care which weights it loads.
The value isn't cheaper tokens — it's AI eligibility for the workloads where a leak invites regulatory scrutiny.
API access is key-gated. Request a key through the contact on the competition entry — then:
curl https://api.fordmorris.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "stable",
"messages": [{"role": "user",
"content": "Hello from the judges"}]}'
401 Unauthorized. That's the production posture, live.