Skip to main content
Hillock ships with a FastAPI server in api.py that exposes an OpenAI-compatible /v1/chat/completions endpoint. Any tool, UI, or library that speaks the OpenAI chat API — Open-WebUI, AnythingLLM, the official openai Python client, Obsidian plugins, or raw curl — can connect to Hillock without modification. Hillock’s HDC gating and knowledge-graph retrieval run transparently behind the standard API surface; the client never needs to know it is talking to a local memory engine rather than a hosted model.

Starting the API Server

The server binds to all interfaces on port 8000. Open http://localhost:8000 in a browser to confirm it is running — you will see a JSON health response:
Ingest your documents via the CLI (python main.py, then /ingest yourfile.pdf) before starting the API server. The API server loads the knowledge graph from the SQLite database on startup — documents ingested after the server starts are available immediately because they share the same live database, but you need content in the graph for meaningful responses.

Connecting Open-WebUI

Open-WebUI is the most popular self-hosted chat frontend for local models. Point it at Hillock’s API server to use Hillock’s grounded memory as the backend.
1

Install and run Open-WebUI

Follow the Open-WebUI installation guide to get the frontend running. The quickest path with Docker:
Open http://localhost:3000 to reach the UI.
2

Go to Settings → Connections → Add Connection

In Open-WebUI, click your profile icon in the top-right corner, select Settings, then navigate to Connections. Click Add Connection (or the + button next to OpenAI API connections).
3

Set API Base URL

In the Base URL field, enter:
If Hillock is running on a different machine on your local network, replace localhost with that machine’s LAN IP address.
4

Set an API key

Open-WebUI requires a non-empty API key field, but Hillock does not enforce authentication. Enter any string:
5

Select the connection and start chatting

Save the connection. Back in the main chat view, select the new Hillock connection from the model dropdown. Open-WebUI will list hillock as the available model (the name your API server reports). Start a conversation — your queries are now routed through Hillock’s memory engine.

Connecting AnythingLLM

AnythingLLM supports custom OpenAI-compatible backends natively.
1

Open AnythingLLM settings

Launch AnythingLLM and click the Settings gear icon, then navigate to AI Providers → LLM.
2

Select Custom OpenAI-compatible provider

From the LLM Provider dropdown, choose Custom OpenAI Compatible.
3

Enter Hillock connection details

Fill in the fields:Save your settings and start a workspace conversation.

Connecting via the OpenAI Python Client

Install the openai package and point it at Hillock’s local server. Your existing OpenAI-based code works without changes beyond the base_url override.

Streaming with the OpenAI Client

Enable SSE streaming by setting stream=True. Hillock’s API server sends the complete response as a single SSE content chunk followed by a finish_reason: stop chunk and [DONE]. The OpenAI client handles this transparently:

Connecting via curl

Test the API directly from your terminal without any client library:
Example response:
For a streaming curl request, add "stream": true to the body:
The server responds with SSE chunks in OpenAI format. The full answer is delivered in a single content chunk, followed by a stop signal:
Obsidian + Smart Connections plugin: If you use Obsidian for note-taking, install the Smart Connections plugin. In its settings, set the API base URL to http://localhost:8000/v1 and the API key to any non-empty string. Smart Connections will route its semantic search and chat queries through Hillock’s grounded memory engine, giving your notes AI-powered Q&A backed by your own ingested documents — with zero data leaving your machine.