2 comments (2 comments)0 reactions (0 reactions)0 assignees (0 assignees)Python514 forks (514 forks)auto 404
help wantednew-task
Repository metrics
- Stars
- 2,496 stars (2,496 stars)
- PR merge metrics
- PR metrics pending (PR metrics pending)
Description
Evaluation short description
Evaluation metadata
Provide all available
Contributor guide
- Research direction
- Study the HELMET benchmark for long context evaluation, understand its data format and metrics, and implement a new task in lighteval that loads and evaluates HELMET long context samples.
- Tech stack
- python
- Domain
- machine learningai
- Issue type
- Feature
- Prerequisites
- PythonGit