Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Llama Cpp

Bookmarks

  1. Running local LLMs with Claude Code via Unsloth

    How to run local LLMs with Claude Code via llama.cpp and Unsloth's GGUF models. The critical fix: Claude Code added an attribution header that breaks KV cache and causes 90% slower inference — must be disabled in settings.json.

All bookmarks