SpaceXAI launched Grok Bot on August 11, 2026, as an early beta for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium.
Each bot runs on a persistent cloud computer with a browser, filesystem, and terminal. It can sign into the apps a person already uses, keep working after the laptop closes, and only surface for approval or a finished job.
That is a harness launch, not a model launch. We do not have a public score for it.
Hermes and OpenClaw have been shipping the open-source version of the same architecture for months: a process that keeps state, memory, and tools across hours or days. Different names, same product category. Value moved from the single response to the outcome that is still correct in the morning.
Our agentic board still times a session.
What SpaceXAI actually shipped
SpaceXAI's launch post and docs are unusually specific. Treat them as vendor claims until someone else reproduces the week, but the architecture is not vague.
| Grok Bot | Hermes | OpenClaw | |
|---|---|---|---|
| Who ships it | SpaceXAI | Nous Research | OpenClaw (Steinberger / community) |
| Runtime | Shared persistent cloud VM | Your machine, VPS, or container | Your hardware, gateway on-device |
| Computer use | Browser, filesystem, terminal; apps without a clean API | Tools, skills, scheduled work | Tools, skills, heartbeat, local files |
| Memory | Named bots keep files, sessions, preferences | Bounded memory across sessions; skills from completed work | Local state and long-running process |
| How you talk to it | Desktop and iOS, like a colleague | CLI plus Telegram, Discord, Slack, WhatsApp | The chat apps you already use |
| Access today | SuperGrok Heavy, Cursor Ultra, Cursor Teams Premium; enterprise waitlist | Self-host, model-agnostic | Self-host, model-agnostic |
| The trade | Low setup, vendor computer, less visibility | Control, you harden it | Control, broader system access |
Two details in the Grok Bot docs matter more than the teammate language.
All of your bots share one computer. Logins, files, and browser sessions are account-scoped, not bot-scoped. SpaceXAI says to treat a login placed on that machine as available to every bot you create. That is a product feature for handoffs. It is also a single blast radius.
Second, the product refuses to begin with a workflow builder. You message a named bot, grant access, and it works in the real tool. If the job is real, the failure modes are also real: a wrong CRM write, a follow-up sent to the wrong thread, a credential sitting on a machine you do not inspect.
We have not run Grok Bot. The table above is the public architecture, not a ranking.
The open-source version was already the product
Hermes is explicit about the loop. It keeps curated memory across sessions, writes skills from completed work, and can sit on a cheap VPS while you reach it from a messaging app. Hermes is selling accumulated competence, not a better first answer.
OpenClaw is the high-adoption expression of the same idea: a local-first assistant that lives in the channels you already use, runs with a heartbeat, and can act without a fresh prompt every time. Flexibility and surface area arrive together. An agent with broader system access can finish more jobs. It can also touch more things you did not intend.
There are other shapes in the same category: collaborative workspaces, coding agents with computer use, enterprise managed agents. Across them, the pattern is stable. Chat optimized for looking intelligent in a turn. These systems optimize for a finished change in someone else's software while you are elsewhere.
If you still evaluate them by reading a chat transcript, you are grading the wrong artifact.
The model was not the missing piece
Tool calling, long context, and multi-step planning got better. That helped. It was not sufficient.
A persistent harness is now the decisive layer: state, a computer that can navigate real interfaces, memory that does not evaporate, scheduled execution, and explicit approval gates. Without those, a frontier model is still a very smart autocomplete that forgets the project between sessions.
Most "agentic" features bolted onto chat products add a few tool calls and stop there. Real agency looks different. The system keeps a running picture of the job. It notices a follow-up is due. It can recover from a partial failure without starting over. It escalates only the decisions that need a human.
That is why launch-day demos often disappoint in week two. Next-token quality is no longer the hard problem. Reliability over long horizons is. So is a website that changed its layout, a credential that expired, and a human who did not notice the agent was quietly doing the wrong thing.
We already see a smaller version of this on coding-agent products. In AI coding agents need receipts we refused to publish a ranking from vendor lists, because the lists rank their authors and the one public cross-agent table still cannot tell you what a failed Tuesday afternoon costs. Persistent general agents are that problem with email, CRM, and billing attached.
What a Terminal-Bench score can still tell you
Our weighted agentic rows are Terminal-Bench 2.0, BrowseComp, and OSWorld-Verified. As of August 2026, the current agentic leader is Claude Opus 5 (80.1, Supported).
Read that sentence for what it is. Those tests ask whether a model can finish a bounded, instrumented task: a terminal workflow, a web-research question, a computer-use job. They are the right shortlist for the brain you put inside a harness. They are a weak score for the harness itself.
| What the public tests measure | What Grok Bot, Hermes, and OpenClaw sell |
|---|---|
| One task, one environment, a clock | Work that continues after you close the laptop |
| A hidden acceptance check at the end | A CRM row, inbox, or ticket that has to stay right overnight |
| The model, sometimes with a standard harness | Memory, credentials, schedules, multi-bot handoff |
| Failures you can see in the log | Failures that look like ordinary business activity |
High Terminal-Bench scores are still useful. Pair them with SWE-bench if the work is software, and with the latency note if the loop has many calls. Then stop. Do not promote an agentic model rank into a verdict on Grok Bot versus Hermes versus OpenClaw. That comparison needs the same jobs, the same success criteria, counted interventions, and published silent failures.
We do not have that fixture for general persistent agents. Until we do, there is no honest product podium.
Who should run the computer
Grok Bot optimizes for low friction and always-on continuity. You do not provision the VM. You do not patch the runtime. Parallel bots keep working on a machine that does not sleep when you do. You trade visibility, pricing, and control of the shared credential boundary for that convenience.
Hermes and OpenClaw optimize for ownership. Your hardware or your VPS. Your choice of model. Data can stay local. You can inspect and extend the agent. The cost is setup, maintenance, and security work that the vendor would otherwise absorb.
Neither side has a permanent advantage. Split them by the work, not by tribe.
Use a managed computer when the work is high-volume, the data is not the crown jewels, and you will actually review the exceptions. Use a self-hosted agent when the logins cannot live on a vendor VM, or when you need to change the harness itself. Many people will end up with both: a managed bot for the always-on grind, a local agent for the sensitive lane.
What matters more than the brand is whether the system can own an outcome across multiple steps without constant supervision, and whether you can see and correct the failures. Black-box "it finished the job" is not a control if the job can move money or customers.
One workflow, two weeks of review
Pick one repetitive job that already costs you real time. Write down what "done" looks like, what must never happen, and how you will check the work. Give a single agent that job and nothing else.
Review every output for two weeks. Instrument the failures: where it got stuck, what it misread, what context was missing. Do not start with five agents and a vague ambition. Compounding starts after the first workflow is trustworthy.
Demand a trajectory you can inspect: tool calls, intermediate state, and the points where a human had to judge. If the product cannot show you that, you are not managing an agent. You are hoping.
Model rankings still help you choose the brain. They do not tell you whether the full system stays reliable over a week of real work. That evaluation is still mostly manual. It will not stay that way, but pretending otherwise is how you ship a wrong Salesforce update and call it agency.
Chatbot windows made sense when models were primarily good at generating text. They no longer match the product people are buying.
The next number worth publishing is a week-horizon eval: the same jobs, counted interventions, silent wrong writes, and a credential scope you can state in one sentence. Until someone runs that across products, we will keep ranking models and refuse to rank the week.
→ Agentic leaderboard · Terminal-Bench 2.0 · OSWorld-Verified · Coding agents need receipts
Reader questions
Frequently asked questions
01What is Grok Bot?
Grok Bot is SpaceXAI's always-on agent product, launched in early beta on August 11, 2026. Each bot runs on a persistent cloud computer with a browser, filesystem, and terminal. It can sign into real apps, keep working after your laptop closes, and only surface for approval or a finished job.
02What is the difference between Grok Bot, Hermes, and OpenClaw?
All three sell a persistent harness, not a chat window. Grok Bot is the managed cloud computer, currently gated to SuperGrok Heavy and certain Cursor plans. Hermes is Nous Research's self-hosted agent with memory and skill creation. OpenClaw is the local-first, high-adoption alternative you run yourself.
03Can BenchLM rank persistent agents?
Not yet. The weighted agentic board scores models on Terminal-Bench, BrowseComp, and OSWorld sessions. Those tests do not measure overnight reliability, silent wrong writes, or credential scope. We will not publish a product ranking of Grok Bot, Hermes, or OpenClaw until the same jobs are run across them.
04Do agentic benchmarks measure always-on agents?
They measure whether a model can finish a bounded, instrumented task with tools. That is necessary for a persistent agent and not sufficient. A high Terminal-Bench score does not tell you whether the system notices a missed follow-up, recovers after a layout change, or updates the wrong Salesforce record at 2 a.m.
05Should I use a managed or self-hosted persistent agent?
Use a managed computer when you want always-on continuity and will accept vendor pricing and less visibility. Use Hermes or OpenClaw when the data and credentials cannot leave your machine. Neither choice removes review. Start with one narrow workflow and audit it for two weeks.
Source ledger
External sources linked in this article
Continue with live BenchLM data
Share or save