Home / Blog / Local LLM stack
Local AI and Backend Architecture

Local LLM Architecture with Ollama, ChromaDB and FastAPI

Running a model on your own machine can improve data control, enable offline work and make experimentation affordable. It does not automatically create a secure or production-ready AI service. A useful architecture separates the HTTP application, inference runtime, embedding model, vector store and source-document lifecycle, then makes latency, memory, authorization and failure behavior visible.

Give each component one clear responsibility

Keep model execution, retrieval and the client-facing contract independently operable
1 / ClientAuthenticated request and bounded input.
2 / FastAPIValidation, policy, timeout and response contract.
3 / RetrievalChroma query with tenant and source filters.
4 / OllamaEmbedding and generation on managed hardware.
5 / ResponseGrounded answer, citations and trace ID.

Separate indexing from question answering

Document ingestion is a controlled job: parse files, attach stable source/version metadata, chunk according to structure, generate embeddings and upsert deterministic IDs. Query-time work should authenticate the caller, enforce access filters, embed the question with the compatible model, retrieve a bounded set of passages and assemble context with source references. Keeping these paths separate makes re-indexing, retries and quality checks easier to manage.

Use the same embedding model and preprocessing for index and query. Changing the model or vector dimensions is a migration: create or rebuild a compatible collection, compare retrieval results and switch deliberately. Store model names, versions, chunking settings and source hashes alongside the corpus metadata.

Design the FastAPI boundary as a real API

BoundaryOperational choiceWhy it matters
InputRequest size, prompt length and allowed file typesBounds memory, latency and abuse
ConcurrencyInference semaphore or job queuePrevents bursts from exhausting RAM/VRAM
TimeoutSeparate retrieval and generation budgetsReturns predictable failures instead of hanging clients
IdentityAuthenticate before tenant-aware vector filtersPrevents cross-user document exposure
ObservabilityTrace ID, model, timings and status without prompt secretsSupports diagnosis without logging sensitive content

Use application lifespan hooks for shared clients and clean shutdown. Keep the API process from loading a separate copy of the model per worker; model memory is often the scarce resource, not Python request handling. Validate deployment behavior under the exact worker and GPU configuration you intend to run.

Plan memory, model loading and throughput

Model weights, context length, batch size, concurrent generations and vector search all compete for resources. Quantization can reduce memory at a quality tradeoff; measure task quality and tokens per second on target hardware rather than assuming a model tag fits. Cold model load, queue wait and generation time are different latency components. Expose them separately in internal metrics.

Bound concurrency and queue length. If callers can wait, return a job identifier and status endpoint rather than holding unlimited HTTP connections. Define what happens when the model is unavailable, disk is full or an index is rebuilding. A local service still needs backups for source files and vector metadata, plus a tested restore path.

Local does not mean private by default

Bind inference and database ports to loopback or a protected private interface; do not expose them directly to the public internet. Put authentication, TLS termination, authorization and rate limits at a deliberate boundary. Keep API keys and source credentials outside the repository, and restrict which users or processes can access the model endpoint and stored documents. Review telemetry, crash reports, temporary uploads and logs because sensitive text can escape through those paths too.

Retrieved files can contain malicious instructions. Treat them as untrusted context, keep tools disabled unless explicitly required, and never let a retrieved passage grant authority. Add deletion and retention workflows that remove source data and derived chunks consistently.

Evaluate the complete service

Build a fixed evaluation set with expected sources, answer requirements and refusal cases. Test retrieval recall and ranking separately from answer grounding, citation accuracy and usefulness. Include concurrent requests, long inputs, empty indexes, stale versions, access boundaries and model restarts. Compare model/quantization changes using both quality and operational metrics.

In summary

Ollama, ChromaDB and FastAPI can form a practical local AI stack when their responsibilities are separated: ingestion owns the corpus, retrieval enforces scope, inference is resource-limited and the API exposes a dependable contract. “Local” is a deployment choice, not a security guarantee. Measure quality, capacity and recovery on the machine and workflow you will actually support.

References