Monday, January 29, 2024
LLM Structured Output with Local Haystack RAG and Ollama
Haystack 2.0 provides functionality to process LLM output and ensure proper JSON structure, based on predefined Pydantic class. I show how you can run this on your local machine, with Ollama. This is possible thanks to OllamaGenerator class available from Haystack.
Tuesday, January 23, 2024
JSON Output with Notus Local LLM [LlamaIndex, Ollama, Weaviate]
In this video, I show how to get JSON output from Notus LLM running locally with Ollama. JSON output is generated with LlamaIndex using the dynamic Pydantic class approach.
Labels:
LlamaIndex,
LLM,
RAG
Monday, January 15, 2024
FastAPI and LlamaIndex RAG: Creating Efficient APIs
FastAPI works great with LlamaIndex RAG. In this video, I show how to build a POST endpoint to execute inference requests for LlamaIndex. RAG implementation is done as part of Sparrow data extraction solution. I show how FastAPI can handle multiple concurrent requests to initiate RAG pipeline. I'm using Ollama to execute LLM calls as part of the pipeline. Ollama processes requests sequentially. It means Ollama will process API requests in the queue order. Hopefully, in the future, Ollama will support concurrent requests.
Labels:
FastAPI,
LlamaIndex,
LLM,
RAG
Monday, January 8, 2024
Transforming Invoice Data into JSON: Local LLM with LlamaIndex & Pydantic
This is Sparrow, our open-source solution for document processing with local LLMs. I'm running local Starling LLM with Ollama. I explain how to get structured JSON output with LlamaIndex and dynamic Pydantic class. This helps to implement the use case of data extraction from invoice documents. The solution runs on the local machine, thanks to Ollama. I'm using a MacBook Air M1 with 8GB RAM.
Labels:
JSON,
LlamaIndex,
LLM,
Pydantic,
RAG
Sunday, December 17, 2023
From Text to Vectors: Leveraging Weaviate for local RAG Implementation with LlamaIndex
Weaviate provides vector storage and plays an important part in RAG implementation. I'm using local embeddings from the Sentence Transformers library to create vectors for text-based PDF invoices and store them in Weaviate. I explain how integration is done with LlamaIndex to manage data ingest and LLM inference pipeline.
Monday, December 11, 2023
Enhancing RAG: LlamaIndex and Ollama for On-Premise Data Extraction
LlamaIndex is an excellent choice for RAG implementation. It provides a perfect API to work with different data sources and extract data. LlamaIndex provides API for Ollama integration. This means we can easily use LlamaIndex with on-premise LLMs through Ollama. I explain a sample app where LlamaIndex works with Ollama to extract data from PDF invoices.
Tuesday, December 5, 2023
Secure and Private: On-Premise Invoice Processing with LangChain and Ollama RAG
The Ollama desktop tool helps run LLMs locally on your machine. This tutorial explains how I implemented a pipeline with LangChain and Ollama for on-premise invoice processing. Running LLM on-premise provides many advantages in terms of security and privacy. Ollama works similarly to Docker; you can think of it as Docker for LLMs. You can pull and run multiple LLMs. This allows to switch between LLMs without changing RAG pipeline.
Subscribe to:
Posts (Atom)