Filesystem RAG benchmarks: corpus/, train.json, evaluate_rag.py (RAGAS quality). Not for prod monitoring, latency/throughput benchmarking (use rag-perf), or evals outside this repo layout.
---
name: rag-eval
version: "2.6.0"
description: >-
Filesystem RAG benchmarks: corpus/, train.json, evaluate_rag.py (RAGAS quality). Not for prod
monitoring, latency/throughput benchmarking (use rag-perf), or evals outside this repo layout.
license: Apache-2.0
compatibility: Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/eval (eval deps live in scripts/eval/pyproject.toml); network to RAG, ingestor, and vdb endpoints; NVIDIA_API_KEY for RAGAS; optional RAG_EVAL_JUDGE_MODEL (default mistralai/mixtral-8x22b-instruct-v0.1).
metadata:
author: NVIDIA RAG <foundational-rag-dev@exchange.nvidia.com>
github-url: "https://github.com/NVIDIA-AI-Blueprints/rag"
endpoint-openapi-schemas:
- docs/api_reference/openapi_schema_rag_server.json
- docs/api_reference/openapi_schema_ingestor_server.json
argument-hint: RAGAS eval | evaluate_rag | train.json | corpus | results json | error triage | uv run --project scripts/eval | enable_reranker | query_rewriting | temperature | skip_ingestion
tags:
- nvidia
- blueprint
- rag
- evaluation
- ragas
- benchmarking
- nvidia-rag-blueprint
languages:
- python… load the full skill through Skill MeIn any Claude conversation, say:
Install the Rag Eval skill
If full content is available, it applies in this conversation and stays installed for future sessions.
Not connected yet? Connect your AI first →
MCP endpoint
https://skillme.dev/api/mcp