rag-eval
無料Filesystem RAG benchmarks: corpus/, train.json, evaluate_rag.py (RAGAS quality). Not for prod monitoring, latency/throughput benchmarking (use rag-perf), or evals outside this repo layout.
日本語の概要は準備中です。原文の説明を表示しています。
7 件 ・ 関連度順
概要と使いどころ
Filesystem RAG benchmarks: corpus/, train.json, evaluate_rag.py (RAGAS quality). Not for prod monitoring, latency/throughput benchmarking (use rag-perf), or evals outside this repo layout.
日本語の概要は準備中です。原文の説明を表示しています。
Implement comprehensive observability for LLM applications including tracing (Langfuse/Helicone), cost tracking, token optimization, RAG evaluation metrics (RAGAS), hallucination detection, and production monitoring. Essential for debugging, optimizing costs, and ensuring AI output quality. Use when ", llm-monitoring, tracing, langfuse, helicone, cost-tracking, ragas, evaluation, hallucination-detection, prompt-caching" mentioned.
日本語の概要は準備中です。原文の説明を表示しています。
Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config. Not for accuracy / RAGAS scoring (use rag-eval) or for deploying / repairing services (use rag-blueprint).
日本語の概要は準備中です。原文の説明を表示しています。
Design RAG pipelines: chunking, retrieval evaluation, and architecture. Use when building a RAG system, selecting a chunking strategy, choosing a vector database, optimizing retrieval quality, or evaluating with RAGAS metrics.
日本語の概要は準備中です。原文の説明を表示しています。
Decomposed, multi-criteria metric design for LLM pipelines. The metric IS the model — change the metric and the optimizer changes behavior. Decompose by default; bool during compile, float during eval; calibrate against human; mitigate judge bias. Search keywords: LLM-as-judge, llm as judge, eval metric, evaluation score, scoring function, rubric, RAGAS, G-Eval, judge bias, verbosity bias, how to evaluate LLM output.
日本語の概要は準備中です。原文の説明を表示しています。
Think and work like an expert Machine Learning Engineer. Use when a task calls for Machine Learning Engineer judgment. Reasons from decision-policy framing, Google's Rules of ML, point-in-time data, and prefill/decode inference physics through GBDT/PyTorch baselines, vLLM/SGLang/KServe serving with FP8/AWQ quantization, hybrid-retrieval RAG, RAGAS and human-validated LLM judges, OpenTelemetry GenAI tracing, and post-Omnibus EU AI Act obligations while treating train-serve skew, temporal leakage, prompt injection, LLM nondeterminism, and degenerate feedback loops as first-class failure modes.
日本語の概要は準備中です。原文の説明を表示しています。
模型评估全流程。覆盖分类/回归/排序/生成/RAG/LLM 评估、A/B 测试、 在线监控、数据漂移检测、评估集构建、基准测试。 触发关键词:评估、指标、Precision、Recall、F1、AUC、BLEU、ROUGE、 A/B 测试、基准测试、数据漂移、模型监控、RAGAS、LLM-as-Judge。
日本語の概要は準備中です。原文の説明を表示しています。