We attempted 48 rewrites across our first-party baseline, Copy.ai, PromptOptimizer.tools, and SiteGPT, then ran every usable prompt three times on the same target model. The study found no prompt-quality winner. On every shared noncoding case, the tools tied on fully accepted outputs.
The observed difference was reach: usable prompt artifacts appeared for 8/12 baseline cases, 6/12 Copy.ai cases, 3/12 PromptOptimizer.tools cases, and 1/12 SiteGPT case. Those are one-session availability observations, not permanent uptime claims.
We picked the tools before seeing their rewrites
DataForSEO selected public tool candidates from the live US desktop results for ai prompt optimizer and prompt improver on July 26, 2026. We screened every eligible result before collecting the first scored rewrite, then froze a three-competitor cap beside the first-party baseline.
| Candidate | Screen decision | Observed boundary |
|---|---|---|
| PromptOptimizer.tools | Selected | Anonymous rewrite and copyable result |
| Copy.ai | Selected | Public form with a required goal field |
| SiteGPT | Selected | Anonymous form with its visible default method |
| Aiven | Eligible, outside cap | Passed after the three places were filled |
| MaxAI | Unavailable | Sign-in required; no authorized session |
| TripleTen | Unavailable | No functional prompt control in available browsers |
| CustomGPT | Unavailable | No usable prompt control appeared |
| PromptPerfect | Unavailable | Signup required; shutdown notice present |
| RightBlogger | Unavailable | Signup required before the prompt field |
No selected competitor was replaced after its outputs became visible. We did not create a paid account, bypass authentication, evade a rate limit, or repeat a failed collection silently.
The study also kept prompt generators separate. A generator turns a brief into a new prompt; an optimizer rewrites an existing one. Scoring both as the same product job would make the table tidy and the conclusion wrong.
Availability changed the usable outcome
Each tool received the same 12 public synthetic cases in a seeded, rotating order. The collector made one attempt per case and preserved all visible failures.
| Tool | Usable artifacts | Unavailable attempts | Observed failure states |
|---|---|---|---|
| First-party baseline | 8/12 | 4/12 | Submit disabled |
| Copy.ai | 6/12 | 6/12 | Input controls unavailable |
| PromptOptimizer.tools | 3/12 | 9/12 | Rate limit and interface errors |
| SiteGPT | 1/12 | 11/12 | No result and interface errors |
These rows describe the interfaces we observed during one collection session. A signed-in plan, later session, different browser, or product update may behave differently.
An optimizer you cannot reach optimizes nothing.
Availability still belongs in an outcome study for that reason. The fair response to a dead interface is not a quality zero. It is to show the missing denominator beside every rate.
Shared cases produced no quality winner
The 18 usable prompt artifacts received random identifiers before target execution. The target Worker saw an artifact ID, rewritten prompt, and public source input, but no tool identity. Every artifact ran three times through openai/gpt-5.6-luna with reasoning disabled, a 4,000-token output limit, zero data retention required, and provider data collection denied.
All 54 target calls completed without a provider error. Automatic checks froze before the mapping from artifact IDs to tools was opened.
| Shared comparison | Cases | Baseline full outputs | Competitor full outputs | Baseline checks | Competitor checks |
|---|---|---|---|---|---|
| Baseline vs Copy.ai | 6 | 8/18 | 8/18 | 59/75 | 57/75 |
| Baseline vs PromptOptimizer.tools | 2 | 3/6 | 3/6 | 21/24 | 17/24 |
| Baseline vs SiteGPT | 1 | 0/3 | 0/3 | 9/12 | 9/12 |
The baseline and Copy.ai tied on fully accepted outputs in each of their six shared noncoding cases. The baseline and PromptOptimizer.tools also tied in both shared noncoding cases. One SiteGPT overlap is too small to distinguish anything.
Conditional full-pass rates across each tool's available cases were 13/24 for the baseline, 8/18 for Copy.ai, 6/9 for PromptOptimizer.tools, and 0/3 for SiteGPT. Those percentages use different task sets and denominators. Presenting them as a leaderboard would reward tools for which cases happened to return a rewrite.
We refuse that ranking.
The work product exposed what prompt polish hid
The provider-policy and research-evidence cases remained partial across several tools because the target answer missed exact citation or table requirements. The prompts looked more complete than the rough request, but the finished work still failed its contract.
That is why this study ran every rewrite through a target model. Prompt length, visual structure, and a successful copy action are intermediate signals. The customer outcome is a cited comparison, valid JSON object, grounded routing decision, or correctly shaped draft that survives its checks.
The shared-case ties do not mean all prompts were identical. No blinded human review judged prose clarity, requirement preservation by inspection, tone, or persuasiveness. The automatic checks asked whether the target output satisfied declared facts and fields. A later human study could find differences this protocol intentionally leaves unclaimed.
The JSON-export coding case also stays outside the usable-work conclusion. PromptOptimizer.tools produced 3/3 full text outputs and the baseline produced 2/3, but neither result proves that a repository patch passed test, lint, and build commands. The separate coding Prompt Lab uses real patches and command logs for that reason.
Inspect the study, including its failed preflight
The first target preflight put the local runner and Wrangler in separate process boundaries. All 54 local fetches failed before reaching the Worker. There were no provider calls, outputs, evaluations, or cost. We preserved that failed artifact instead of quietly deleting it: the corrected run changed only the process placement, while the collected prompts, blinding procedure, target settings, and automatic checks stayed fixed.
The successful 54 target runs cost $0.179673. Baseline prompts accounted for $0.054093 of that target-model cost, Copy.ai for $0.035724, PromptOptimizer.tools for $0.070995, and SiteGPT for $0.018861. Different output lengths drove much of the spread, and this one route is not a price forecast.
The public files expose the decisions and failures needed to audit the result:
- Preregistered study manifest
- Access-screen decisions
- All 48 public-tool attempts
- All 54 blinded target runs
- Preserved failed preflight
npm run validate:prompt-tool-study recalculates the cohort, artifact counts, target runs, blinding boundary, and automatic evaluations without making a provider request.
The AI Prompt Optimizer is the baseline this study tested; the AI Prompt Generator handles the separate blank-brief job. Either one starts a test — neither is proof that the finished work is ready.
What the comparison can support
This study supports a narrow promise: optimize prompts to reduce unusable outputs and rework, then verify that claim on representative tasks.
It does not support “best prompt optimizer.” The session covered 12 synthetic cases, one target-model route, three runs per available artifact, and public interfaces as they appeared on one date. Four candidates were unavailable, one eligible candidate sat outside the frozen cap, and no blinded human quality review ran.
The next comparison should add more shared cases and a preregistered blinded rubric. Until then, availability and accepted work stay in separate columns.
Reader questions
Frequently asked questions
01Which AI prompt optimizer performed best?
This study did not find a prompt-quality winner. On every noncoding case shared by the first-party baseline and a tested competitor, the tools tied on fully accepted target outputs. The baseline returned usable prompt artifacts in more of the twelve attempted cases, but unequal availability prevents that reach result from becoming an overall quality ranking.
02How were the prompt optimizers compared?
Each selected tool received the same twelve public synthetic cases through its offered interface. Every usable rewritten prompt received a random identifier and ran three times on the same target-model route. Frozen automatic checks judged the finished work. Failed interfaces and rate limits remained unavailable rather than receiving quality zeros.
03Why does prompt-optimizer availability matter?
A rewrite that never appears cannot improve the target work. Availability was measured as one session observation, not permanent uptime: the baseline returned eight usable artifacts, Copy.ai six, PromptOptimizer.tools three, and SiteGPT one. Those denominators remain visible beside all conditional quality rates.
04Did the study test ChatGPT prompt generators?
No. A prompt optimizer begins with an existing prompt and returns a rewrite. A prompt generator begins with a brief and creates a new prompt. Combining those jobs would reward or penalize tools for solving different problems, so generator candidates were recorded for separate research and excluded from this comparison.
05Can these results predict future prompt-optimizer performance?
No. This was one session on July 26, 2026, using twelve synthetic cases, public free interfaces, and one target-model route. Interfaces, rate limits, upstream models, and provider behavior can change. The artifacts support a reproducible snapshot and a better evaluation method, not a permanent market ranking.
Continue with live BenchLM data
Share or save