Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@
*.jl.mem
/Manifest*.toml
/docs/Manifest*.toml
/test/Manifest*.toml
/docs/build/
.env
.env.example
Expand Down
6 changes: 2 additions & 4 deletions Project.toml
Original file line number Diff line number Diff line change
Expand Up @@ -9,11 +9,10 @@ HuggingFaceHub = "d0076355-e2c0-48e6-a044-05906e51b7fc"
JSON3 = "0f8b85d8-7281-11e9-16c2-39a750bddbf1"
LibPQ = "194296ae-ab2e-5f79-8cd4-7183a0a5a0d1"
LinearAlgebra = "37e2e46d-f89d-539d-b4ee-838fcccc9c8e"
OpenAI = "e9f21f70-7185-4079-aca2-91159181367c"
PromptingTools = "670122d1-24a8-4d70-bfce-740807c42192"
RAGTools = "16ddad29-bbe8-45a7-857d-3d9514eb0023"
Serialization = "9e88b42a-f829-5b0c-bbe9-9e923198166b"
SparseArrays = "2f01184e-e22b-5df5-ae63-d93ebab69eaf"
Statistics = "10745b16-79ce-11e8-11f9-7d13ad32a3b2"
URIs = "5c2747f8-b7ea-4ff2-ba2e-563bfd36b1d4"

[compat]
Expand All @@ -22,10 +21,9 @@ HuggingFaceHub = "0.1.2"
JSON3 = "1.14.3"
LibPQ = "1.18.0"
LinearAlgebra = "1.10"
OpenAI = "0.11"
PromptingTools = "0.82.1"
RAGTools = "0.7.0"
Serialization = "1.10"
SparseArrays = "1.10"
Statistics = "1.10"
URIs = "1.6"
julia = "1.10"
9 changes: 5 additions & 4 deletions docs/make.jl
Original file line number Diff line number Diff line change
Expand Up @@ -8,19 +8,20 @@ makedocs(;
authors="ParamThakkar123 <paramthakkar864@gmail.com> and TheCedarPrince <jacobszelko@gmail.com>",
sitename="HealthLLM.jl",
format=Documenter.HTML(;
canonical="https://ParamThakkar123.github.io/HealthLLM.jl",
edit_link="master",
canonical="https://JuliaHealth.github.io/HealthLLM.jl",
edit_link="main",
assets=String[],
),
pages=[
"Home" => "index.md",
"Getting Started" => "getting-started.md",
"Document Ingestion" => "ingestion.md",
"Building Embeddings" => "embeddings.md",
"Querying the RAG System" => "querying.md",
],
)

deploydocs(;
repo="github.com/ParamThakkar123/HealthLLM.jl",
devbranch="master",
repo="github.com/JuliaHealth/HealthLLM.jl",
devbranch="main",
)
76 changes: 70 additions & 6 deletions docs/src/embeddings.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,9 +58,73 @@ embedding_ref("all-minilm"; provider=:huggingface) # "hf:sentence-transf

!!! note "Backend setup"
For Ollama, pull the tag once with `ollama pull nomic-embed-text` and make
sure the server is running. For HuggingFace, ensure the corresponding schema
is configured in `PromptingTools`. Either way, register the model in the RAG
pipeline with `register_models(...)` as shown in [Getting Started](getting-started.md).
sure the server is running. For HuggingFace, set a token (see below). Either
way, register the model in the RAG pipeline with `register_models(...)` as
shown in [Getting Started](getting-started.md).

### The HuggingFace backend

PromptingTools ships schemas for a long list of OpenAI-compatible providers but
none for HuggingFace, so this package supplies one:
[`HuggingFaceOpenAISchema`](@ref). It is an `AbstractOpenAISchema`, so
`aigenerate`, `aiembed`, message rendering and retries all work unchanged —
`provider = :huggingface` and `hf:`-prefixed model names route through it
automatically.

Set a token once per session, or let it come from the environment
(`HF_API_TOKEN`, `HF_TOKEN`, `HUGGINGFACE_API_KEY`, `HUGGING_FACE_HUB_TOKEN`,
checked in that order):

```julia
set_huggingface_api_key!(ENV["HF_TOKEN"])
huggingface_api_key() # what will actually be sent
```

!!! note "Pinning an inference provider"
The router auto-routes only to providers **enabled on your account**, so a
model can be live on HuggingFace and still be refused with
`model_not_supported`. Append the provider to pin it:

```julia
huggingface_providers("Qwen/Qwen2.5-7B-Instruct") # ["featherless-ai"]

register_models("hf:Qwen/Qwen2.5-7B-Instruct:featherless-ai", "hf:BAAI/bge-m3")
```

When routing is refused, the error raised here already names the live
providers and the exact model string to use, so you rarely need to look this
up yourself. Alternatively enable the provider at
[huggingface.co/settings/inference-providers](https://huggingface.co/settings/inference-providers).

Chat and embeddings use different HuggingFace surfaces, because the router's
OpenAI-compatible API covers chat completions only — `/v1/embeddings` answers
404 there:

| Call | Endpoint |
|--------------|---------------------------------------------------------------------------|
| `aigenerate` | [`HUGGINGFACE_ROUTER_URL`](@ref) + `/chat/completions` |
| `aiembed` | [`HUGGINGFACE_INFERENCE_URL`](@ref) + `/<model>/pipeline/feature-extraction` |

The feature-extraction reply is reshaped into the OpenAI embeddings response
`aiembed` expects, and token-level output (from models that do not pool
internally) is mean-pooled to one vector per input — so a HuggingFace embedding
matrix is the same `dim × n` shape as an Ollama one.

!!! note "Cold models"
A HuggingFace model that is not already warm loads while holding the
connection open — around 40–55s for `bge-m3` in practice. That exceeds
PromptingTools' 120s `aiembed` default under load, so this package raises the
read timeout to [`HUGGINGFACE_EMBED_TIMEOUT`](@ref) (300s) when you have not
chosen one yourself. Any `http_kwargs` you pass is left exactly as given.

To use a deployment that *does* speak OpenAI embeddings — a Text Embeddings
Inference container or a dedicated Inference Endpoint — pass its URL and the
request is forwarded there instead:

```julia
E = embed(chunks, "bge-m3"; provider = :huggingface,
api_kwargs = (; url = "https://my-endpoint.hf.space/v1"))
```

## Generating embeddings

Expand Down Expand Up @@ -186,7 +250,7 @@ conn = LibPQ.Connection("postgresql://user:pass@localhost/health")
store = PgVectorStore(conn, embedding_dimension(); table = "omop_embeddings", metric = :cosine)

add!(store, E, chunks) # creates the table + inserts in one transaction
hits = search(store, embed("count patients"), 5) # (; id, chunk, distance)
hits = search(store, embed("count patients"), 5) # Vector{Hit}: row id in `index`, raw `distance` kept
```

`metric` chooses the distance operator: `:cosine` (`<=>`), `:dot` (`<#>`), or
Expand All @@ -196,7 +260,7 @@ wrapper:

```julia
store_embeddings_pgvector(conn, E, chunks, embedding_dimension(); table = "omop_embeddings")
hits = search_embeddings_pgvector(conn, embed("count patients"), 5; table = "omop_embeddings")
hits = search_embeddings_pgvector(conn, embed("count patients"), 5; table = "omop_embeddings") # raw (; id, chunk, distance)
```

!!! note "pgvector prerequisites"
Expand All @@ -217,7 +281,7 @@ using Faiss # optional; load before con

store = FaissVectorStore(embedding_dimension()) # inner-product index (cosine after normalisation)
add!(store, E, chunks)
hits = search(store, embed("count patients"), 5) # (; index, chunk, score)
hits = search(store, embed("count patients"), 5) # Vector{Hit}, as with every backend
```

As with the local store, vectors are normalised for the inner-product/cosine
Expand Down
5 changes: 4 additions & 1 deletion docs/src/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,10 @@ corpus = write_combined_file(files, "corpus.txt")

### 2. Register models

Register chat and embedding models for use with `PromptingTools`. Supports Ollama, HuggingFace, and other backends:
Register chat and embedding models for use with `PromptingTools`. Model names
prefixed with `hf:` are routed to HuggingFace via [`HuggingFaceOpenAISchema`](@ref)
(set a token with [`set_huggingface_api_key!`](@ref) or `HF_TOKEN`); everything
else defaults to Ollama. See [Building Embeddings](embeddings.md) for the details:

```julia
register_models("llama3.2", "nomic-embed-text")
Expand Down
3 changes: 2 additions & 1 deletion docs/src/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ The package centers on five areas:
- collecting source files and writing combined corpora
- ingesting curated docs and web-search results into an index (see [Document Ingestion](ingestion.md))
- building retrieval indexes through `RAGTools`
- constructing grounded prompts from retrieved chunks (see [Querying the RAG System](querying.md))
- generating retrieval-backed answers for query construction
- building and validating embeddings across Ollama and HuggingFace (see [Building Embeddings](embeddings.md))
- storing embeddings in a local file, PostgreSQL/`pgvector`, or FAISS
Expand All @@ -32,5 +33,5 @@ index = build_index_rag(RAGTools.SimpleIndexer(), files)
More detailed setup, testing commands, and the end-to-end walkthrough are in [Getting Started](getting-started.md).

```@autodocs
Modules = [HealthLLM, HealthLLM.Utils, HealthLLM.Database, HealthLLM.Query, HealthLLM.Ingestion, HealthLLM.Embeddings, HealthLLM.Storage]
Modules = [HealthLLM, HealthLLM.HuggingFace, HealthLLM.Utils, HealthLLM.Database, HealthLLM.Prompt, HealthLLM.Execution, HealthLLM.Query, HealthLLM.Ingestion, HealthLLM.Embeddings, HealthLLM.Storage]
```
Loading
Loading