Skip to main content

Celeris

Data verified

Celeris Labs is an AI research company building low-latency language models. Its first model, Celeris-1, uses diffusion-based decoding and is served through an OpenAI-compatible API.

Model count
1
Best model
N/A
Celeris-1
Avg. top 3
N/A
Latest release
July 2026
Public-ranked
0
Supported
0

Provider profile

What is Celeris?

Celeris Labs is an artificial intelligence research company focused on reducing language-model inference time. Tom Hamer and Jesse Clark founded the company after previously founding Marqo. Celeris says its research spans model training, diffusion-based architecture, and the production systems required to serve language models at interactive speeds.

Celeris-1 is the company’s first public model. It is a proprietary, general-purpose diffusion language model released on July 22, 2026. The hosted API uses the familiar chat-completions shape, supports streaming, and exposes an 8,192-token context window. Current public access is limited to the United States.

Celeris at a glance

Company
Celeris Labs
Founders
Tom Hamer and Jesse Clark
First model
Celeris-1
Architecture
Diffusion language model
API
OpenAI-compatible chat completions
Public availability
United States

What the published evidence supports

The exact benchmark row we can attach today is Celeris-1’s 75.9% MMLU-Pro result. Celeris ran the full test set with five-shot chain-of-thought examples, strict answer-format scoring, and the reasoning budget set to zero. The result is provider-published rather than independently reproduced, so it remains visible with its source while the model stays outside the overall ranking.

Celeris separately reports 158 ms p50 response time on the MMLU-Pro run and 1,664 output tokens per second at p50 on a 1,000-token workload. Those methods are public, but they are not the same measurement pipeline used by our runtime table. We therefore keep the claims in the source record and leave the independent speed field blank.

Where Celeris-1 fits

The model is aimed at short, structured, latency-sensitive work: routing, extraction, classification, query rewriting, agent steps, and interfaces that cannot hide a long generation behind a loading state. OpenAI compatibility lowers the integration cost because existing SDK calls can point at Celeris with a base-URL change.

The limit is just as concrete. Celeris-1 has an 8K context window, and the official model guide recommends another model for long-form generation. One provider-run knowledge benchmark also cannot establish coding, agentic, multilingual, or instruction-following performance. Test the actual request shape before putting it on a critical path.

Provider models

Browse all models
ContextPricePublic score
Celeris-1Current8K$2.00 / $6.00N/AN/A

Celeris FAQ

What is Celeris AI?

Celeris Labs is an AI research company focused on low-latency language-model training, architecture, and inference. Its first public model is Celeris-1, a proprietary diffusion language model served through an OpenAI-compatible API. The company was founded by Tom Hamer and Jesse Clark, who previously founded Marqo.

What model does Celeris offer?

Celeris currently lists Celeris-1 as its public model. It has an 8,192-token context window and is designed for short, structured requests where response time matters. The provider recommends another model for long-form generation, and public API availability is currently limited to the United States.

Is Celeris-1 a diffusion LLM?

Yes. Celeris describes Celeris-1 as a diffusion language model that can decode multiple tokens per model invocation instead of relying only on sequential, token-by-token generation. The company has not released the model weights or a complete architecture specification, so that description remains a provider claim rather than an independently inspected implementation.

How much does the Celeris-1 API cost?

Celeris lists pay-as-you-go pricing at $2 per million input tokens and $6 per million output tokens, metered separately. The plan includes streaming and the OpenAI-compatible API. Enterprise customers can ask for volume pricing, dedicated clusters, custom limits, VPC deployment, and a service-level agreement.

How fast is Celeris-1?

Celeris reports 1,664 output tokens per second at p50 on its long-form speed workload and 158 ms p50 response time on its MMLU-Pro run. Both are provider-run measurements with published methods. They are not inserted into BenchLM’s independent runtime field because the measurement pipelines differ.

Compare Celeris models