TextGen
TextGen is an open-source interface for running language models on your own hardware. The repository describes portable desktop builds for Windows, Linux and macOS, with GPU and CPU-only options. The portable builds support GGUF models through llama.cpp; additional backends and training features use a fuller installation.
Match the installation to the model
For a portable setup, download a compatible GGUF model and place it in the user_data/models folder. Multi-file Transformers and EXL3 models need their own subfolders and the full installation. Choosing the right format is more useful than assuming any downloaded model will load in every build.
The repository lists CUDA, Vulkan, ROCm and CPU-only choices. Select a build for your hardware and check model memory requirements before downloading a large checkpoint. Local software removes a hosted provider from the generation path, but model files, hardware and electricity still have costs.
The fuller installation adds backends such as Transformers and ExLlamaV3, plus LoRA training, image generation and extensions. Its requirements are heavier than the portable chat app. Install only the capabilities needed for your intended work.
Use the interface and API with compatible models
TextGen offers instruction-following chat, character-oriented modes and a notebook for free-form generation. You can edit messages, branch a conversation and attach text files, PDFs or Word documents. Vision input depends on a suitable multimodal model and its configuration.
The repository advertises OpenAI- and Anthropic-compatible APIs, including chat, completions and messages endpoints. It also supports tool calling through Python tools and MCP servers. Compatibility describes an interface, not ownership by those providers or equal output quality. Test the particular client, request shape and model before switching an application.
Custom tools can fetch web pages or execute functions. A local inference setup that invokes those tools is no longer an offline-only task. Review tool permissions and the destinations they contact before using confidential data.
Keep the service private by configuration
The project describes no telemetry and an offline local interface. Its command-line options also allow network listening, public sharing and a public API. Those are deployment choices that change who can reach the service.
Before exposing an endpoint, configure authentication and restrict access. The documented API key and separate admin key distinguish inference access from operations such as loading models. Keep both credentials out of public client code.
TextGen uses AGPL-3.0. Review the license for distribution or service use, and review model licenses separately. Begin with a small model and a representative prompt, then inspect latency and memory on your actual machine. API compatibility and quantization settings do not establish a universal speed or quality guarantee.
