Official repo for the ACSOS 2021 paper on how to manage many deep learning models at the edge!
-
Updated
Sep 29, 2021 - Python
Official repo for the ACSOS 2021 paper on how to manage many deep learning models at the edge!
Local inference serving with adaptive batching, benchmark sweeps, and regression gates for batching tradeoffs.
Reactive LLM inference.
cost-aware-inference-cluster: FastAPI inference-serving prototype with Redis queues, worker heartbeats, dynamic batching, autoscaling simulation, and local Docker Compose benchmark evidence.
GSM artifact for exact online GNN inference serving with degree-based scheduling and memory-aware batch division.
GPU-aware LLM serving runtime with async scheduling, dynamic batching, streaming, admission control, metrics, and mock/llama.cpp/vLLM backends.
Machine-readable companion to the IEEE OJ-CS survey 'Semantic Caching and Response Reuse for Large Language Model Services: A Survey' (Chukkapalli, Mishra, Naik, 2026): 21-work evidence matrix, systematic-search log, proposed benchmark trace schema, stdlib-only contract validator, and CPU pilot. Code MIT; data CC-BY-4.0.
Add a description, image, and links to the inference-serving topic page so that developers can more easily learn about it.
To associate your repository with the inference-serving topic, visit your repo's landing page and select "manage topics."