huggingface/lighteval
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
https://skillcdn.ai/gh/huggingface/lightevalConnect this address to your AI to use these skills. How to connect
- Unverified
- Default branch
main - Commit
c2af9e5 - License
MIT
Explore this repository
Searches the original text of skills and documents; display translations are not searched.
Folders
Documents
- _toctree.ymldocs/source/_toctree.yml
- Adding a Custom TaskLighteval provides a flexible framework for creating custom evaluation tasks. This guide explains how to create and integrate new tasks into the evaluation system.docs/source/adding-a-custom-task.mdx
- Adding a New MetricThere are two types of metrics in Lighteval:docs/source/adding-a-new-metric.mdx
- Available tasksBrowse and inspect tasks available in LightEval.docs/source/available-tasks.mdx
- Caching SystemLighteval includes a caching system that can significantly speed up evaluations by storing and reusing model predictions. This is especially useful when running the same evaluation multiple times, or…docs/source/caching.mdx
- Contributing to Multilingual EvaluationsLighteval supports multilingual evaluations through a comprehensive system of translation literals and language-adapted templates.docs/source/contributing-to-multilingual-evaluations.mdx
- Evaluating Custom ModelsLighteval allows you to evaluate custom model implementations by creating a custom model class that inherits from LightevalModel. This is useful when you want to evaluate models that aren't directly…docs/source/evaluating-a-custom-model.mdx
- Lighteval🤗 Lighteval is your all-in-one toolkit for evaluating Large Language Models (LLMs) across multiple backends with ease. Dive deep into your model's performance by saving and exploring detailed, sampl…docs/source/index.mdx
- Evaluate your model with Inspect-AIPick the right benchmarks with our benchmark finder: Search by language, task type, dataset name, or keywords.docs/source/inspect-ai.mdx
- InstallationLighteval can be installed from PyPI or from source. This guide covers all installation options and dependencies.docs/source/installation.mdx
- Metric ListThese metrics use log-likelihood of the different possible targets.docs/source/metric-list.mdx
- Offline evaluation using local data filesIf you are prototyping a task based on files that are not yet hosted on the Hub, you can take advantage of the hf_data_files argument to point lighteval at local JSON/CSV resources. This makes it eas…docs/source/offline-evaluation.md
- Quick TourLighteval can be used with several different commands, each optimized for different evaluation scenarios.docs/source/quicktour.mdx
- Saving and Reading ResultsLighteval provides comprehensive logging and result management through the EvaluationTracker class. This system allows you to save results locally and optionally push them to various platforms for co…docs/source/saving-and-reading-results.mdx
- Using Hugging Face Inference Endpoints or TGI as BackendAn alternative to launching the evaluation locally is to serve the model on a TGI-compatible server/container and then run the evaluation by sending requests to the server. The command is the same as…docs/source/use-huggingface-inference-endpoints-or-tgi-as-backend.mdx
- Using Inference Providers as BackendLighteval allows you to use Hugging Face's Inference Providers to evaluate LLMs on supported providers such as Black Forest Labs, Cerebras, Fireworks AI, Nebius, Together AI, and many more.docs/source/use-inference-providers-as-backend.mdx
- Using LiteLLM as BackendLighteval allows you to use LiteLLM as a backend, enabling you to call all LLM APIs using the OpenAI format. LiteLLM supports various providers including Bedrock, Hugging Face, Vertex AI, Together AI…docs/source/use-litellm-as-backend.mdx
- Using SGLang as BackendLighteval allows you to use SGLang as a backend, providing significant speedups for model evaluation. To use SGLang, simply change the model_args to reflect the arguments you want to pass to SGLang.docs/source/use-sglang-as-backend.mdx
- Using VLLM as BackendLighteval allows you to use VLLM as a backend, providing significant speedups for model evaluation. To use VLLM, simply change the model_args to reflect the arguments you want to pass to VLLM.docs/source/use-vllm-as-backend.mdx
- Using the Python APILighteval can be used from a custom Python script. To evaluate a model, you will need to set up an [~logging.evaluation_tracker.EvaluationTracker], [~pipeline.PipelineParameters], a model or a model_…docs/source/using-the-python-api.mdx