api.py that exposes an OpenAI-compatible /v1/chat/completions endpoint. Any tool, UI, or library that speaks the OpenAI chat API — Open-WebUI, AnythingLLM, the official openai Python client, Obsidian plugins, or raw curl — can connect to Hillock without modification. Hillock’s HDC gating and knowledge-graph retrieval run transparently behind the standard API surface; the client never needs to know it is talking to a local memory engine rather than a hosted model.
Starting the API Server
8000. Open http://localhost:8000 in a browser to confirm it is running — you will see a JSON health response:
Ingest your documents via the CLI (
python main.py, then /ingest yourfile.pdf) before starting the API server. The API server loads the knowledge graph from the SQLite database on startup — documents ingested after the server starts are available immediately because they share the same live database, but you need content in the graph for meaningful responses.Connecting Open-WebUI
Open-WebUI is the most popular self-hosted chat frontend for local models. Point it at Hillock’s API server to use Hillock’s grounded memory as the backend.1
Install and run Open-WebUI
Follow the Open-WebUI installation guide to get the frontend running. The quickest path with Docker:Open
http://localhost:3000 to reach the UI.2
Go to Settings → Connections → Add Connection
In Open-WebUI, click your profile icon in the top-right corner, select Settings, then navigate to Connections. Click Add Connection (or the
+ button next to OpenAI API connections).3
Set API Base URL
In the Base URL field, enter:If Hillock is running on a different machine on your local network, replace
localhost with that machine’s LAN IP address.4
Set an API key
Open-WebUI requires a non-empty API key field, but Hillock does not enforce authentication. Enter any string:
5
Select the connection and start chatting
Save the connection. Back in the main chat view, select the new Hillock connection from the model dropdown. Open-WebUI will list
hillock as the available model (the name your API server reports). Start a conversation — your queries are now routed through Hillock’s memory engine.Connecting AnythingLLM
AnythingLLM supports custom OpenAI-compatible backends natively.1
Open AnythingLLM settings
Launch AnythingLLM and click the Settings gear icon, then navigate to AI Providers → LLM.
2
Select Custom OpenAI-compatible provider
From the LLM Provider dropdown, choose Custom OpenAI Compatible.
3
Enter Hillock connection details
Fill in the fields:
Save your settings and start a workspace conversation.
Connecting via the OpenAI Python Client
Install theopenai package and point it at Hillock’s local server. Your existing OpenAI-based code works without changes beyond the base_url override.
Streaming with the OpenAI Client
Enable SSE streaming by settingstream=True. Hillock’s API server sends the complete response as a single SSE content chunk followed by a finish_reason: stop chunk and [DONE]. The OpenAI client handles this transparently:
Connecting via curl
Test the API directly from your terminal without any client library:"stream": true to the body: