Quantizes LLMs to 8-bit or 4-bit for 50-75% memory reduction with minimal accuracy loss. Use when GPU memory is limited, need to fit larger models, or want faster inference. Supports INT8, NF4, FP4 fo
10-optimization/bitsandbytes/SKILL.md
Use this skill when the user asks to save, remember, recall, or organize memories. Triggers on: 'remember this', 'save t...
CLI tool for configuring and monitoring Claude Code
指导Claude按照二哥的风格撰写求职类文章,包括公司薪资爆料、年终奖盘点、求职攻略、offer选择建议等内容。