Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GP
devops/serving-llms-vllm
Creates GitHub pull requests with properly formatted titles that pass the check-pr-title CI validation. Use when creatin...
Create a database migration to add a table, add columns to an existing table, add a setting, or otherwise change the sch...
A set of resources to help me write all kinds of internal communications, using the formats that my company likes to use...