Auteur • 198 livres
PII Redaction with Presidio : Privacy‑Safe Logs, Prompts, and Documents
Trex Team
Llama Guard in Practice : Safety Classification Pipelines for LLM Responses
RRF in Practice : Hybrid Search That Beats Pure Vector Retrieval
Jina AI Search Stack : Embeddings, Rerankers, and Hybrid Retrieval End‑to‑End
SPLADE Sparse Retrieval : Modern BM25‑Style Search for RAG Pipelines
Production GraphRAG : Building Knowledge Graph Retrieval for Verifiable Enterprise Answers
BGE Rerankers : Training and Deploying Modern Cross‑Encoder Reranking
Guardrails AI : Validating Outputs with Schemas, Rules, and Safe Retries
ColBERT Reranking : High‑Recall Retrieval Without Blowing Your Latency Budget
Vespa for AI Search : Serving Hybrid Retrieval at Web‑Scale
Weaviate for RAG : Schema‑Aware Vector Search with Hybrid Queries
Meilisearch for Developers : Fast, Typo‑Tolerant Search for Modern Apps
Typesense Search : Instant Search APIs with Relevance You Can Tune
LanceDB : Columnar Vector Storage for Fast Local and Cloud Retrieval
Milvus at Scale : Vector Indexing, Sharding, and Retrieval Operations
Qdrant Vector Search : Shipping High‑Recall Retrieval with Filters and Payloads
Tantivy in Rust : Building High‑Performance Full‑Text Search Services
Chroma in Production : Lightweight Vector Stores for Prototypes That Grow Up
pgvector for Postgres : Practical Vector Search Without a New Database
MongoDB Atlas Vector Search : Retrieval‑Ready Schemas for Operational Data
Redis Vector Similarity : Real‑Time Retrieval and Caching for AI Apps
Elasticsearch ELSER : Semantic Search Without Managing Embeddings
vLLM Serving : High‑Throughput LLM APIs with PagedAttention and KV Cache Tuning
OpenSearch Neural Search : Hybrid Retrieval, Reranking, and Operations
llama.cpp in the Real World : On‑Device LLM Inference, Profiling, and Deployment
Ollama for Teams : Local Model Distribution, Versioning, and Secure Developer Workflows
TensorRT‑LLM Optimization : Quantization, Kernel Fusion, and Throughput Engineering
Text Generation Inference (TGI) : Deploying Transformers with Streaming and Batching
SGLang : Structured Generation for Tool Use, JSON Outputs, and Fast Inference
BentoML : Packaging, Deploying, and Monitoring ML/LLM Services End‑to‑End
KServe on Kubernetes : Production Model Serving with Canary Releases and Autoscaling
Ray Serve for LLM Apps : Scalable APIs, Batching, and Async Tool Pipelines
NVIDIA Triton Inference Server in Practice : Operating Multi‑Model GPU Inference at Scale
NVIDIA NIM in Practice : Standardized Inference Microservices for Enterprise GenAI
bitsandbytes : Practical 8‑bit/4‑bit Optimization for LLM Training and Inference
GGUF Model Packaging : Managing Quantized LLM Artifacts for Local Inference
36 de 198 titres