Skip to main content
BenchLM Researchprompt optimizer comparison

Four Prompt Optimizers, One Test Set

We attempted 48 public prompt-optimizer rewrites and ran every usable artifact three times. Shared cases tied on accepted outputs; availability was the observed difference.

Published
Last updated
Reading time
8 min
External sources
0
Tags: prompt optimizer comparison, AI prompt optimizer, prompt improver, prompt optimizationData and scoring methodology
In this article6 sections

We attempted 48 rewrites across our first-party baseline, Copy.ai, PromptOptimizer.tools, and SiteGPT, then ran every usable prompt three times on the same target model. The study found no prompt-quality winner. On every shared noncoding case, the tools tied on fully accepted outputs.

The observed difference was reach: usable prompt artifacts appeared for 8/12 baseline cases, 6/12 Copy.ai cases, 3/12 PromptOptimizer.tools cases, and 1/12 SiteGPT case. Those are one-session availability observations, not permanent uptime claims.

We picked the tools before seeing their rewrites

DataForSEO selected public tool candidates from the live US desktop results for ai prompt optimizer and prompt improver on July 26, 2026. We screened every eligible result before collecting the first scored rewrite, then froze a three-competitor cap beside the first-party baseline.

Table 1
Candidate Screen decision Observed boundary
PromptOptimizer.tools Selected Anonymous rewrite and copyable result
Copy.ai Selected Public form with a required goal field
SiteGPT Selected Anonymous form with its visible default method
Aiven Eligible, outside cap Passed after the three places were filled
MaxAI Unavailable Sign-in required; no authorized session
TripleTen Unavailable No functional prompt control in available browsers
CustomGPT Unavailable No usable prompt control appeared
PromptPerfect Unavailable Signup required; shutdown notice present
RightBlogger Unavailable Signup required before the prompt field

No selected competitor was replaced after its outputs became visible. We did not create a paid account, bypass authentication, evade a rate limit, or repeat a failed collection silently.

The study also kept prompt generators separate. A generator turns a brief into a new prompt; an optimizer rewrites an existing one. Scoring both as the same product job would make the table tidy and the conclusion wrong.

Availability changed the usable outcome

Each tool received the same 12 public synthetic cases in a seeded, rotating order. The collector made one attempt per case and preserved all visible failures.

Across twelve attempted cases, BenchLM returned eight usable prompt artifacts, Copy.ai six, PromptOptimizer.tools three, and SiteGPT one.

Table 2
Tool Usable artifacts Unavailable attempts Observed failure states
First-party baseline 8/12 4/12 Submit disabled
Copy.ai 6/12 6/12 Input controls unavailable
PromptOptimizer.tools 3/12 9/12 Rate limit and interface errors
SiteGPT 1/12 11/12 No result and interface errors

These rows describe the interfaces we observed during one collection session. A signed-in plan, later session, different browser, or product update may behave differently.

An optimizer you cannot reach optimizes nothing.

Availability still belongs in an outcome study for that reason. The fair response to a dead interface is not a quality zero. It is to show the missing denominator beside every rate.

Shared cases produced no quality winner

The 18 usable prompt artifacts received random identifiers before target execution. The target Worker saw an artifact ID, rewritten prompt, and public source input, but no tool identity. Every artifact ran three times through openai/gpt-5.6-luna with reasoning disabled, a 4,000-token output limit, zero data retention required, and provider data collection denied.

All 54 target calls completed without a provider error. Automatic checks froze before the mapping from artifact IDs to tools was opened.

Table 1
Shared comparison Cases Baseline full outputs Competitor full outputs Baseline checks Competitor checks
Baseline vs Copy.ai 6 8/18 8/18 59/75 57/75
Baseline vs PromptOptimizer.tools 2 3/6 3/6 21/24 17/24
Baseline vs SiteGPT 1 0/3 0/3 9/12 9/12

The baseline and Copy.ai tied on fully accepted outputs in each of their six shared noncoding cases. The baseline and PromptOptimizer.tools also tied in both shared noncoding cases. One SiteGPT overlap is too small to distinguish anything.

Conditional full-pass rates across each tool's available cases were 13/24 for the baseline, 8/18 for Copy.ai, 6/9 for PromptOptimizer.tools, and 0/3 for SiteGPT. Those percentages use different task sets and denominators. Presenting them as a leaderboard would reward tools for which cases happened to return a rewrite.

We refuse that ranking.

The work product exposed what prompt polish hid

The provider-policy and research-evidence cases remained partial across several tools because the target answer missed exact citation or table requirements. The prompts looked more complete than the rough request, but the finished work still failed its contract.

That is why this study ran every rewrite through a target model. Prompt length, visual structure, and a successful copy action are intermediate signals. The customer outcome is a cited comparison, valid JSON object, grounded routing decision, or correctly shaped draft that survives its checks.

The shared-case ties do not mean all prompts were identical. No blinded human review judged prose clarity, requirement preservation by inspection, tone, or persuasiveness. The automatic checks asked whether the target output satisfied declared facts and fields. A later human study could find differences this protocol intentionally leaves unclaimed.

The JSON-export coding case also stays outside the usable-work conclusion. PromptOptimizer.tools produced 3/3 full text outputs and the baseline produced 2/3, but neither result proves that a repository patch passed test, lint, and build commands. The separate coding Prompt Lab uses real patches and command logs for that reason.

Inspect the study, including its failed preflight

The first target preflight put the local runner and Wrangler in separate process boundaries. All 54 local fetches failed before reaching the Worker. There were no provider calls, outputs, evaluations, or cost. We preserved that failed artifact instead of quietly deleting it: the corrected run changed only the process placement, while the collected prompts, blinding procedure, target settings, and automatic checks stayed fixed.

The successful 54 target runs cost $0.179673. Baseline prompts accounted for $0.054093 of that target-model cost, Copy.ai for $0.035724, PromptOptimizer.tools for $0.070995, and SiteGPT for $0.018861. Different output lengths drove much of the spread, and this one route is not a price forecast.

The public files expose the decisions and failures needed to audit the result:

npm run validate:prompt-tool-study recalculates the cohort, artifact counts, target runs, blinding boundary, and automatic evaluations without making a provider request.

The AI Prompt Optimizer is the baseline this study tested; the AI Prompt Generator handles the separate blank-brief job. Either one starts a test — neither is proof that the finished work is ready.

What the comparison can support

This study supports a narrow promise: optimize prompts to reduce unusable outputs and rework, then verify that claim on representative tasks.

It does not support “best prompt optimizer.” The session covered 12 synthetic cases, one target-model route, three runs per available artifact, and public interfaces as they appeared on one date. Four candidates were unavailable, one eligible candidate sat outside the frozen cap, and no blinded human quality review ran.

The next comparison should add more shared cases and a preregistered blinded rubric. Until then, availability and accepted work stay in separate columns.

Reader questions

Frequently asked questions

01Which AI prompt optimizer performed best?

This study did not find a prompt-quality winner. On every noncoding case shared by the first-party baseline and a tested competitor, the tools tied on fully accepted target outputs. The baseline returned usable prompt artifacts in more of the twelve attempted cases, but unequal availability prevents that reach result from becoming an overall quality ranking.

02How were the prompt optimizers compared?

Each selected tool received the same twelve public synthetic cases through its offered interface. Every usable rewritten prompt received a random identifier and ran three times on the same target-model route. Frozen automatic checks judged the finished work. Failed interfaces and rate limits remained unavailable rather than receiving quality zeros.

03Why does prompt-optimizer availability matter?

A rewrite that never appears cannot improve the target work. Availability was measured as one session observation, not permanent uptime: the baseline returned eight usable artifacts, Copy.ai six, PromptOptimizer.tools three, and SiteGPT one. Those denominators remain visible beside all conditional quality rates.

04Did the study test ChatGPT prompt generators?

No. A prompt optimizer begins with an existing prompt and returns a rewrite. A prompt generator begins with a brief and creates a new prompt. Combining those jobs would reward or penalize tools for solving different problems, so generator candidates were recorded for separate research and excluded from this comparison.

05Can these results predict future prompt-optimizer performance?

No. This was one session on July 26, 2026, using twelve synthetic cases, public free interfaces, and one target-model route. Interfaces, rate limits, upstream models, and provider behavior can change. The artifacts support a reproducible snapshot and a better evaluation method, not a permanent market ranking.

Share or save

Share on XShare on LinkedIn

Keep reading

All research

New models drop every week. Join 2,000+ readers for one email a week on what moved, why, and what still needs proof.