# FLTEval

> A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.

Canonical page: https://benchlm.ai/benchmarks/flteval

- Category: [Coding](/coding)
- Last updated: September 27, 2026

## About FLTEval

- Year: 2026
- Tasks: FLT project pull requests
- Format: Lean 4 repository task completion
- Difficulty: Formal verification / proof engineering
- Paper: [Leanstral: Open-Source foundation for trustworthy vibe-coding](https://mistral.ai/news/leanstral)

FLTEval is designed to move evaluation beyond isolated competition-math problems. Instead of proving one-off statements, models must operate inside realistic formal repositories and finish pull-request-style Lean 4 work with Lean itself acting as a verifier.

FLTEval is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does FLTEval measure?

A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.

### Which model scores highest on FLTEval?

No models have been evaluated on FLTEval yet.

### How many models are evaluated on FLTEval?

0 AI models have been evaluated on FLTEval on BenchLM.

### Does FLTEval affect BenchLM's overall score?

Not directly. FLTEval is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
