> ## Documentation Index
> Fetch the complete documentation index at: https://hillock.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Connect Hillock to Open-WebUI, AnythingLLM, and Local UIs

> Start Hillock's FastAPI server and connect any OpenAI-compatible frontend — Open-WebUI, AnythingLLM, Obsidian, or curl — in under five minutes.

Hillock ships with a FastAPI server in `api.py` that exposes an OpenAI-compatible `/v1/chat/completions` endpoint. Any tool, UI, or library that speaks the OpenAI chat API — Open-WebUI, AnythingLLM, the official `openai` Python client, Obsidian plugins, or raw `curl` — can connect to Hillock without modification. Hillock's HDC gating and knowledge-graph retrieval run transparently behind the standard API surface; the client never needs to know it is talking to a local memory engine rather than a hosted model.

## Starting the API Server

```bash theme={null}
python api.py
# 🚀 Starting Hillock OpenAI-Compatible API Server on http://0.0.0.0:8000
```

The server binds to all interfaces on port `8000`. Open `http://localhost:8000` in a browser to confirm it is running — you will see a JSON health response:

```json theme={null}
{
  "status": "online",
  "message": "Hillock API is running! Point your AI client to /v1/chat/completions",
  "version": "0.6.1"
}
```

<Note>
  Ingest your documents via the CLI (`python main.py`, then `/ingest yourfile.pdf`) **before** starting the API server. The API server loads the knowledge graph from the SQLite database on startup — documents ingested after the server starts are available immediately because they share the same live database, but you need content in the graph for meaningful responses.
</Note>

## Connecting Open-WebUI

Open-WebUI is the most popular self-hosted chat frontend for local models. Point it at Hillock's API server to use Hillock's grounded memory as the backend.

<Steps>
  <Step title="Install and run Open-WebUI">
    Follow the [Open-WebUI installation guide](https://docs.openwebui.com) to get the frontend running. The quickest path with Docker:

    ```bash theme={null}
    docker run -d -p 3000:8080 \
      -v open-webui:/app/backend/data \
      --name open-webui \
      ghcr.io/open-webui/open-webui:main
    ```

    Open `http://localhost:3000` to reach the UI.
  </Step>

  <Step title="Go to Settings → Connections → Add Connection">
    In Open-WebUI, click your profile icon in the top-right corner, select **Settings**, then navigate to **Connections**. Click **Add Connection** (or the `+` button next to OpenAI API connections).
  </Step>

  <Step title="Set API Base URL">
    In the Base URL field, enter:

    ```
    http://localhost:8000/v1
    ```

    If Hillock is running on a different machine on your local network, replace `localhost` with that machine's LAN IP address.
  </Step>

  <Step title="Set an API key">
    Open-WebUI requires a non-empty API key field, but Hillock does not enforce authentication. Enter any string:

    ```
    hillock-local
    ```
  </Step>

  <Step title="Select the connection and start chatting">
    Save the connection. Back in the main chat view, select the new Hillock connection from the model dropdown. Open-WebUI will list `hillock` as the available model (the name your API server reports). Start a conversation — your queries are now routed through Hillock's memory engine.
  </Step>
</Steps>

## Connecting AnythingLLM

AnythingLLM supports custom OpenAI-compatible backends natively.

<Steps>
  <Step title="Open AnythingLLM settings">
    Launch AnythingLLM and click the **Settings** gear icon, then navigate to **AI Providers → LLM**.
  </Step>

  <Step title="Select Custom OpenAI-compatible provider">
    From the LLM Provider dropdown, choose **Custom OpenAI Compatible**.
  </Step>

  <Step title="Enter Hillock connection details">
    Fill in the fields:

    | Field | Value |
    | - | - |
    | Base URL | `http://localhost:8000/v1` |
    | API Key | any non-empty string (e.g., `hillock-local`) |
    | Chat Model | `hillock` |

    Save your settings and start a workspace conversation.
  </Step>
</Steps>

## Connecting via the OpenAI Python Client

Install the `openai` package and point it at Hillock's local server. Your existing OpenAI-based code works without changes beyond the `base_url` override.

```python theme={null}
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="hillock-local"  # any non-empty string
)

response = client.chat.completions.create(
    model="hillock",
    messages=[{"role": "user", "content": "What did Alan Turing crack?"}]
)

print(response.choices[0].message.content)
# Alan Turing cracked the Enigma cipher during World War II...
```

### Streaming with the OpenAI Client

Enable SSE streaming by setting `stream=True`. Hillock's API server sends the complete response as a single SSE content chunk followed by a `finish_reason: stop` chunk and `[DONE]`. The OpenAI client handles this transparently:

```python theme={null}
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="hillock-local"
)

stream = client.chat.completions.create(
    model="hillock",
    messages=[{"role": "user", "content": "Where was Marie Curie born?"}],
    stream=True
)

for chunk in stream:
    delta = chunk.choices[0].delta
    if delta.content:
        print(delta.content, end="", flush=True)

print()  # Final newline
```

## Connecting via curl

Test the API directly from your terminal without any client library:

```bash theme={null}
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "hillock", "messages": [{"role": "user", "content": "What is Paris the capital of?"}]}'
```

Example response:

```json theme={null}
{
  "id": "chatcmpl-3f8a1b2c...",
  "object": "chat.completion",
  "created": 1718200000,
  "model": "hillock",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Paris is the capital of France."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {"prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0}
}
```

For a streaming curl request, add `"stream": true` to the body:

```bash theme={null}
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "hillock", "messages": [{"role": "user", "content": "Who discovered radioactivity?"}], "stream": true}'
```

The server responds with SSE chunks in OpenAI format. The full answer is delivered in a single content chunk, followed by a stop signal:

```
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Marie Curie discovered radioactivity."},"finish_reason":null}]}

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]
```

<Tip>
  **Obsidian + Smart Connections plugin:** If you use Obsidian for note-taking, install the [Smart Connections](https://github.com/brianpetro/obsidian-smart-connections) plugin. In its settings, set the API base URL to `http://localhost:8000/v1` and the API key to any non-empty string. Smart Connections will route its semantic search and chat queries through Hillock's grounded memory engine, giving your notes AI-powered Q\&A backed by your own ingested documents — with zero data leaving your machine.
</Tip>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.