Skip to main content
Ryan Orban

Ryan Orban

Subject
2 entries

Local Llm

Bookmarks

  1. Running local LLMs with Claude Code via Unsloth

    How to run local LLMs with Claude Code via llama.cpp and Unsloth's GGUF models. The critical fix: Claude Code added an attribution header that breaks KV cache and causes 90% slower inference — must be disabled in settings.json.

  2. Dalai: Run LLaMA Locally with One Command

    Dalai is a one-command installer for running LLaMA models locally — npm install to set up, then query models via CLI or socket server. One of the first tools to make local LLM inference accessible to developers without ML expertise.

All bookmarks