Category

Local AI

Ollama LM Studio Open WebUI vLLM llama.cpp SGLang TensorRT-LLM Text Generation Inference (TGI) Local RAG GPU Optimization

31 posts

Maximize Local LLM Batch Sizes

Deploying large language models locally on consumer-grade GPUs presents a unique set of challenges. While high-end data center accelerators offer vast memory pools, the average enthusiast is often constrained by 8GB to 24GB of VRAM. This limitation frequently forces developers to choose between a...

Mastering Text Generation Inference: High-Performance Local LLM Deployment

As the Large Language Model (LLM) landscape evolves, the gap between cutting-edge research models and production-ready inference has narrowed significantly. For developers and data scientists, the challenge is no longer just accessing models, but serving them efficiently, securely, and at scale. ...

Build Local RAG Pipeline with LlamaIndex

In today's data-driven landscape, enterprises are increasingly hesitant to send sensitive internal documents to public Large Language Model (LLM) APIs. The solution? A fully local Retrieval-Augmented Generation (RAG) pipeline. By combining LlamaIndex for data structuring and ChromaDB for vector s...