cryptopoly / ChaosEngineAI Sponsor Star 24 Code Issues Pull requests Local AI workstation — discover, run, chat, benchmark, and generate images from open-weight models. DFlash/DDTree speculative decoding, TurboQuant & TriAttention cache compression strategies, MLX + llama.cpp + vLLM + MTPLX backends. desktop-app python machine-learning typescript ai image-generation mlx tauri huggingface apple-silicon openai-api cache-compression llm stable-diffusion llama-cpp vllm local-ai gguf speculative-decoding dflash Updated Jul 24, 2026 Python
Labeeb2339 / recurquant Star 0 Code Issues Pull requests Packed INT4/INT8 recurrent-state cache for Qwen3.5 Gated DeltaNet, with reproducible fidelity evaluation. research transformers pytorch quantization mixed-precision linear-attention cache-compression llm-inference qwen3 qwen35 gated-deltanet recurrent-state Updated Jul 26, 2026 Python