Open-source inference server and production cluster for all the models your agent needs.
-
Updated
Jul 25, 2026 - Python
Open-source inference server and production cluster for all the models your agent needs.
SPLADE: sparse neural search (SIGIR21, SIGIR22)
Neural Search
An easy-to-use python toolkit for flexibly adapting various neural ranking models to target domain.
Fast search index for SPLADE sparse retrieval models implemented in Python using Numpy and Numba
Lite weight wrapper for the independent implementation of SPLADE++ models for search & retrieval pipelines. Models and Library created by Prithivi Da, For PRs and Collaboration checkout the readme.
Provides a minimal PyTorch implementation of SPLADE
Optimized RAG Retrieval with Indexing, Quantization, Hybrid Search and Caching
Local-first RAG system with hybrid (dense + BM25) and SPLADE retrieval, hierarchical conversational memory, and document-aware reasoning.
Advanced RAG pipeline with hybrid dense+sparse search (fixed-size, recursive, semantic chunking), cross-encoder reranking, and RAGAS-based evaluation across all config combinations. Streamlit dashboard included.
A production-grade Multi-Agent RAG (Retrieval-Augmented Generation) system designed for scalable, low-latency, and reliable AI-powered retrieval. Built with hybrid search, cross-encoder reranking, intelligent query decomposition, semantic caching, adaptive LLM routing, and ONNX-optimized inference using Qdrant, Groq, Gemini, and BGE embeddings.
GPU-accelerated semantic search for your docs and source code - hybrid dense + sparse RAG on a local Qdrant backend, served to Claude Code and other MCP clients. The search companion to vaultspec-core.
Learning-to-Rank on MS MARCO Passages: candidate generation from prebuilt indexes and re-ranking for QA search
Training Data Generator for SPLADE Model Fine-tuning
Research: signal-processing approaches to semantic search, treating documents as continuous signals
Official code and camera-ready analyses for PFW Task 8 at SemEval-2026 Task 8 (MTRAGEval).
QLoRA fine-tuning pipeline: Llama 3.1 & Qwen 2.5 for structured extraction — SPLADE sparse embeddings, vLLM FP8 inference, end-to-end data pipeline
Memsplora - An in-memory SPLADE (SParse Lexical AnD Expansion) content server with FAISS integration
Add a description, image, and links to the splade topic page so that developers can more easily learn about it.
To associate your repository with the splade topic, visit your repo's landing page and select "manage topics."