TALE: training-free task-aware layer pruning for improving LLM accuracy and reducing inference cost
nlp transformers pruning model-compression efficient-inference layer-pruning task-specialization llms
-
Updated
Jul 22, 2026 - Python