Skip to main content

Ask questions of your own documents with a language model that runs entirely on your machine. Oasis indexes PDFs, text and CSVs into a local Chroma database, retrieves the relevant passages, and answers offline.

Source on GitHub
  • Solo Project
  • Python
  • LangChain + Chroma
  • Hugging Face
From source file to local index
  1. Files.txt · .pdf · .csv
  2. Chunks1,000 characters · 200 overlap
  3. EmbeddingsInstructor-XL
  4. Index on diskChroma

Indexing the files

Put files in the ingestion folder and run the ingestion command. LangChain splits them into overlapping chunks, Instructor-XL embeds the chunks, and Chroma stores the vectors on disk. I took the approach from PrivateGPT.

Answering a question

The terminal menu in main.py starts a query session. Oasis finds nearby chunks in Chroma, passes them to the local model as context, and prints an answer. You can then ask to see the source passages before entering another question.

The default model is Vicuna 7B through a Hugging Face text-generation pipeline. Its temperature is set to zero. Once the models are downloaded, queries use the index on disk and run offline.

One query through Oasis
  1. QuestionEntered in the terminal
  2. RetrieveClosest chunks from Chroma
  3. GenerateLocal model with retrieved context
  4. InspectAnswer and optional source passages

Where it goes next

Ingestion is manual today: changed files are copied in and ingested again. The next step is to automate that, so Oasis can follow a source, crawl a website or keep up with a project’s documentation on its own.

Running the model takes a capable GPU. I wrote and tested Oasis on an RTX 4090 with 24 GB of VRAM, so the plan is to put it behind a chat app, where one machine answers for everyone who can message it. Microsoft Teams comes first, then Slack and Discord. Oasis is MIT-licensed and on GitHub.

Read the source