One clean Go interface to chat LLMs — Anthropic, OpenAI, ChatGPT subscription plans, OpenRouter, and self-hosted servers (vLLM, Ollama, any OpenAI-compatible endpoint) — with normalized streaming, tool calling, structured output, reasoning, usage, cost, and errors.
go-llm is a low-level provider client library: it does one layer well
and stops there. No agent loops, no prompt frameworks, no magic. The core
package has zero third-party dependencies; official vendor SDKs are used
where they exist and pulled in only by the provider package you import.
- One
llm.Providerinterface — blockingChatand streamingChatStream(iter.Seq2iterators), the same request/response model everywhere, per-provider escape hatches down to the raw SDK client. - Providers: Anthropic (API key), OpenAI
(Responses API), OpenAI Codex (ChatGPT Plus/Pro subscription OAuth),
OpenRouter, and self-hosted vLLM. First-party presets run the shared offline
conformance suite, and credentialed presets have capability-driven live
scenarios.
chatcompletions.New(baseURL)covers any other OpenAI-compatible server (Ollama, llama.cpp, Groq, Together, ...). - Subscription auth: OpenAI Codex can mint browser PKCE credentials or
consume and auto-refresh compatible existing credentials.
llm.LoadAuthFilereads the documented credential union; minted and renewed tokens are returned to your code to persist. - Tools: parallel calls, streamed arguments, and a defined contract for
malformed tool calls — rescue what's rescuable, drop the rest visibly
(
ToolCallDropped), with an opt-inllm.RetryDroppedToolCallsmiddleware. OpenAI-compatible no-argument tool schemas are sent withparameters.properties: {}; non-objectpropertiesvalues fail locally. - Structured output:
llm.Parse[T]decodes model output straight into your struct (native JSON-schema mode where supported, forced-tool or JSON-mode fallback elsewhere), andschema.For[T]generates JSON Schema from Go types for tool inputs, withschema.ValidateArgsfor checking model-emitted arguments. - Reasoning: one
Effortdial across providers; reasoning output is normalized, and raw provider payloads (signed thinking blocks, encrypted reasoning items) round-trip so same-provider replay just works. - Persistence & portability: canonical, versioned JSON for messages and responses. A tool-using conversation started on one provider can be serialized and continued on another — cross-provider handoff is part of the live test matrix, not a hope.
- Usage & cost: normalized tokens (including cache reads/writes and
reasoning tokens),
CostUSDfrom native provider reporting (OpenRouter) or estimated from an embedded, refreshable models.dev price table, including request-wide long-context pricing tiers — with provenance inCostSource(nativevsestimated) — plusContextUsagefor context-window accounting. Live model catalogs preserve independent input/output/cache-read/cache-write availability, so unknown rates are not displayed as free while explicit zero rates remain visible. - Capabilities: provider-wide
Capabilities()discovery with pre-flight request validation, plus advisory per-model capabilities from rich live catalogs such as OpenRouter and Codex. Unsupported provider features fail fast withErrUnsupported; incomplete model metadata never blocks a call. - Conveniences:
Session(auto-managed history, session tools withAddToolResults+Continue, cumulative usage),History,PromptTemplate,llm.Ptrfor optional scalar fields, middleware viallm.Wrap, and observability built in but silent by default (sloglogging,UsageTracker, wire capture viaWithWireCapture). Sensitive headers are always redacted; captured URLs and bodies remain application-sensitive data. Built-in failure logs use safe provider-error summaries and never emit provider-controlled bodies or metadata. - Testing:
llmtest— likenet/http/httptest, but for code that consumes go-llm. - CLI:
llm-cli, a curl-like frontend built entirely on the public API.
go get github.com/pkieltyka/go-llmGo 1.26 is the minimum version for users of the module. Development and releases use Go 1.27.0 or newer, while CI retains a Go 1.26 compatibility lane. Provider SDK dependencies are pulled only when importing provider packages.
package main
import (
"context"
"fmt"
llm "github.com/pkieltyka/go-llm"
"github.com/pkieltyka/go-llm/providers/openai"
)
func main() {
ctx := context.Background()
p, err := openai.New() // reads OPENAI_API_KEY
if err != nil {
panic(err)
}
resp, err := p.Chat(ctx, &llm.Request{
Model: "gpt-5.5",
Messages: []llm.Message{llm.UserText("Explain Go iterators in one sentence.")},
})
if err != nil {
panic(err)
}
fmt.Println(resp.Text())
}Streaming:
for text, err := range llm.StreamText(p.ChatStream(ctx, req)) {
if err != nil {
return err
}
fmt.Print(text)
}Structured output:
type Summary struct {
Title string `json:"title"`
Tags []string `json:"tags"`
}
summary, resp, err := llm.Parse[Summary](ctx, p, &llm.Request{
Model: "gpt-5.5",
Messages: []llm.Message{llm.UserText("Summarize this release note.")},
})
_, _ = summary, resp| Package | Auth | Notes |
|---|---|---|
providers/anthropic |
ANTHROPIC_API_KEY or WithAPIKey |
Messages API |
providers/openai |
OPENAI_API_KEY or WithAPIKey |
Responses API — reasoning survives across tool-call turns |
providers/openaicodex |
WithOAuth only (ChatGPT Plus/Pro) |
Responses wire shape; explicit Models calls use cached authenticated discovery with a curated fallback |
providers/openrouter |
OPENROUTER_API_KEY or WithAPIKey |
Chat Completions; routing/plugins, explicit rich model discovery, and native per-request cost reporting |
providers/vllm |
optional WithAPIKey (vLLM --api-key) |
Self-hosted vLLM preset: host-first, era-aware, live-tested |
providers/ollama |
none | Data-only local-Ollama preset over the engine below (community-verified) |
providers/chatcompletions |
optional WithAPIKey |
Public engine for ANY OpenAI-compatible server: New(baseURL, ...) + declarative Compat quirks |
anthropic.New() // ANTHROPIC_API_KEY
openai.New() // OPENAI_API_KEY
openrouter.New() // OPENROUTER_API_KEYOpenAI Responses callers can select the user-visible reasoning-summary detail
without adding an OpenAI-specific field to llm.Request:
resp, err := p.Chat(ctx, &llm.Request{
Model: "gpt-5.5",
Effort: llm.EffortHigh,
Messages: []llm.Message{llm.UserText("Compare the two proposals.")},
ProviderOptions: openai.Options{
ReasoningSummary: openai.ReasoningSummaryConcise,
},
})The supported selectors are ReasoningSummaryAuto,
ReasoningSummaryConcise, and ReasoningSummaryDetailed. Leaving the field
empty preserves the existing behavior (summary: "auto" when Effort is
set); unknown values fail before a request is sent. OpenAI remains
authoritative about which models accept each selector. This option is not
available on the separate openaicodex subscription provider.
OpenAI Codex consumes credentials minted by compatible existing tools, refreshes them automatically, and hands renewals back to your code to persist; the root library never writes credential files itself:
auth, err := llm.LoadAuthFile("auth.json") // documented credential union
if err != nil {
panic(err)
}
codex, err := openaicodex.New(openaicodex.WithOAuth(auth["openai-codex"], func(ctx context.Context, updated llm.AuthCredential) error {
return persistCredential(ctx, updated)
}))Codex can also mint the initial credential through its provider-owned browser
PKCE flow. Begin returns a redacting loopback launch target; Complete
waits for the automatic callback, while Submit accepts a copied complete
callback URL/query if the browser cannot reach localhost. The host must store
the returned credential before passing it to WithOAuth:
flow, err := openaicodex.NewLoginFlow()
authorization, err := flow.Begin(ctx)
showLoginURL(authorization.URL())
credential, err := flow.Complete(ctx)
err = persistCredential(ctx, credential)The callback binds only 127.0.0.1, uses the registered localhost path on
port 1455 with port 1457 as its fallback, and never persists the credential.
Live OAuth verification is still required because automated tests cannot
complete an account login. Device authorization and a built-in CLI login UX
remain deferred.
Persistence callbacks must honor their context and return only after the
rotated credential is durably stored; an error prevents the provider from
publishing it. A credential containing a refresh token requires a non-nil
callback; access-only credentials may pass nil. To deliberately keep
rotations only in memory, pass an explicit context-aware no-op:
discardRotation := func(ctx context.Context, _ llm.AuthCredential) error {
if err := ctx.Err(); err != nil {
return err
}
return nil
}This can leave the stored refresh token stale after restart and should be a
conscious application decision. Provider-specific request extensions live in
each provider's Options type, passed through Request.ProviderOptions; the
raw SDK client is always reachable via each provider's Client(). Ordinary
openai.Options fields use go-llm and standard-library types; only advanced
escape hatches such as Provider.Client, chatcompletions.Dialect,
chatcompletions.Config, and Provider.BuildParams are vendor-coupled and
stability-exempt before v1.
Self-hosted servers are first-class: constructors are host-first and the
API key is optional (no environment fallback). The providers/vllm
preset targets the current stable vLLM v0.26.0 protocol and knows its
dialect — reasoning output (parsed to
llm.ReasoningPart, streamed as ReasoningDelta), Effort →
reasoning_effort, choice-less mid-stream error events, max_model_len as
ModelInfo.ContextWindow, and typed extensions (top_k, min_p,
stop_token_ids, chat_template_kwargs, thinking_token_budget,
vllm_xargs, plus native
structured_outputs constraint modes — regex/choice/grammar/structural-tag
via vllm.StructuredOutputs; JSON schema stays on the unified
ResponseFormat):
p, err := vllm.New("http://localhost:8000/v1") // keyless by default
if err != nil {
panic(err)
}
model, err := p.ResolveModel(ctx, "qwen") // resolves deployment prefixes/suffixes
if err != nil {
panic(err)
}
resp, err := p.Chat(ctx, &llm.Request{
Model: model.ID,
Effort: llm.EffortNone, // thinking-by-default models answer tersely
Messages: []llm.Message{llm.UserText("hello")},
ProviderOptions: vllm.Options{
XArgs: map[string]any{"custom_engine_arg": "1"},
},
})For deployments that support it, vllm.Options.ThinkingTokenBudget caps
reasoning without adding a provider-neutral knob. When MaxTokens is set,
go-llm clamps the configured budget to reserve 1,024 tokens for the visible
answer and rejects thinking-disabled combinations before sending the request.
vLLM also gives you exact token counting: the preset exposes the
server's /tokenize endpoints as typed extension methods, so context
accounting is ground truth (server-rendered chat template, tools included)
instead of an estimate — live-verified to match a real request's
prompt_tokens exactly:
result, err := p.Tokenize(ctx, req) // same conversion + validation as Chat
usage := result.ContextUsage() // exact count vs the model's max_model_len
fmt.Printf("prompt occupies %d of %d tokens (%.2f%%)\n",
usage.UsedTokens, usage.Window, usage.UsedPercent)Detokenize and TokenizerInfo round out the family. For everything else
that speaks the Chat Completions shape, the same engine is public:
// Ollama's local convention, as a data-only preset:
p, err := ollama.New("") // http://localhost:11434/v1
// Or any OpenAI-compatible endpoint, with quirks declared as data:
p, err = chatcompletions.New("https://api.example.com/v1",
chatcompletions.WithName("example"),
chatcompletions.WithAPIKey(os.Getenv("EXAMPLE_API_KEY")),
chatcompletions.WithCompat(chatcompletions.Compat{StreamIncludeUsage: true}),
)Bonus recipe: current vLLM v0.26.0 also serves an
Anthropic /v1/messages endpoint, so the Anthropic provider can target a
vLLM box directly — anthropic.New(anthropic.WithBaseURL("http://localhost:8000"), anthropic.WithAPIKey("dummy")) completes chats with usage and even maps
Qwen thinking output to reasoning parts (vLLM emits Anthropic-style thinking
blocks there).
llm.MarshalMessages / llm.UnmarshalMessages give you a canonical,
versioned JSON envelope for conversation history — safe to store in a
database and reload across releases. Because every adapter re-encodes the
neutral history into its own wire shape (and drops what only the original
provider can accept, like signed thinking blocks), a saved conversation can
be continued on a different provider: mid-conversation failover,
cheap-model tool loops handed to a stronger model for synthesis, or histories
that simply outlive your vendor choice.
go install github.com/pkieltyka/go-llm/cmd/llm-cli@latestllm-cli -p openai -m gpt-5.5 "write a short haiku about Go"
echo "long input" | llm-cli -p anthropic -m claude-opus-4-8 -s "summarize stdin"
llm-cli -p openrouter -m openai/gpt-5.5 --usage --json "return a JSON status"
llm-cli -p openai-codex --auth-file ~/.config/llm/auth.json -m gpt-5.4 "review this diff"
llm-cli models -p openrouter # list models (table or --json)
llm-cli -p openai -m gpt-5.5 --save chat.json "Start a checklist"
llm-cli -p anthropic -m claude-opus-4-8 --load chat.json --save chat.json "Continue it"llm-cli models explicitly performs provider model-discovery network I/O.
OpenRouter uses exactly its /models response and does not fall back to
models.dev. Listings include advisory supported/default reasoning effort,
whether the provider positively reports reasoning as required, and model
capabilities. Text output shows EFFORTS, DEFAULT EFFORT, and
REASONING REQUIRED; --json emits the corresponding optional fields.
ReasoningRequired is selection metadata only and never overrides the
caller's Request.Effort.
The last pair saves a conversation with one provider and continues it with another — the handoff described above, from the shell.
For OpenAI Codex, authentication precedence is --auth-file, then
OPENAI_CODEX_ACCESS_TOKEN, then compatibility --api-key. Prefer the first
two forms: command-line API-key values are visible in process arguments and
often shell history. Auth files are loaded only when explicitly requested.
Refresh publication holds a bounded cross-process advisory lock across the
full read/modify/atomic-write transaction so concurrent CLI processes do not
lose one another's credential updates.
Use llmtest for unit tests that should never contact a provider — like
net/http/httptest, but for code that consumes go-llm:
p := llmtest.New(llmtest.WithCapabilities(llm.CapabilityJSONSchema))
p.EnqueueResponse(&llm.Response{Parts: []llm.Part{llm.Text(`{"ok":true}`)}})Script responses, streams, and errors; assert on the requests your code made
via p.Requests(). See examples/testing for a complete
worked example.
Provider authors should run two complementary offline suites. RunConformance
checks the common lifecycle, streaming, and normalized-result contract.
RunCapabilityConformance checks that each reviewed advertised capability
activates its exact native request fields and returns the expected normalized
result. Its CapabilityInvocation is fixture control data passed to an
isolated provider factory; it never changes the real request or its model.
Explicit profile exemptions mean “offline evidence gap,” not “unsupported.”
Neither suite proves live account/model availability, quota, cache admission,
or service reliability; credentialed e2e scenarios remain the live evidence.
The advanced chatcompletions.NewWithDialect conformance suite intentionally
stays base-only; the public generic engine owns the reusable compatible-engine
activation profile.
Repository commands:
make build # compile all packages and create bin/llm-cli
make test # offline tests with the race detector
make check # complete credential-free, network-dependent CI gate
make models # refresh the embedded root models.json snapshot and provenanceLive end-to-end tests are behind the live build tag and read credentials
from gollm-test.json (copy gollm-test.json.sample; missing credentials
skip visibly, never fail):
make e2e-testEvery program in examples/ is dual-mode: it runs offline
out of the box against the scripted llmtest provider, and switches to the
real API when ANTHROPIC_API_KEY or OPENROUTER_API_KEY is set.
go run ./examples/chat # offline, scripted fallback
ANTHROPIC_API_KEY=... go run ./examples/chat # same program, real API| Example | Shows |
|---|---|
chat |
One blocking request/response |
stream |
Streaming deltas with llm.StreamText |
tools |
A full tool round trip: call → local execution → result → answer |
parse |
llm.Parse[T] structured output into a Go struct |
history-replay |
Serialize a conversation and continue it on another provider (a real Anthropic→OpenRouter handoff when both keys are set) |
observability |
UsageTracker, logging, and cost via llm.Wrap middleware |
provider-selection |
Choosing a provider at runtime: env keys, an LLM_AUTH_FILE credential file (including codex OAuth), or the offline fake |
testing |
Unit-testing your code with llmtest: request assertions, streaming, error paths (go test ./examples/testing/) |
Godoc examples in example_test.go run offline as part of the test suite.
Pre-1.0: public APIs are intended to be small and stable, but breaking
changes may happen before v1.0.0 when provider behavior or the unified
model needs correction (the chatcompletions.Dialect interface is
explicitly an advanced, stability-exempt surface — prefer the declarative
Compat). After v1.0.0, standard Go module compatibility rules apply.
MIT.