Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Build or fix software · Code with a local model

Best local LLM for coding — September 2026

On BenchLM's public evidence, GLM-5.3 has the highest coding score estimate among the models that meet this page's constraints (61.4). Open weights are only the first requirement. Hardware, quantization, license, and tool access need verification.

Shortlist data updated

Share this shortlist

Share on XLinkedIn

This address is permanent. It always shows the current shortlist for this job, with the date the data was last updated. The embed shows the same shortlist on your site and links back here.

The shortlist for local LLM for coding

Constraints on this page: a builder choosing by accuracy, running on their own hardware, ordinary input size. Change any of them under Refine. A small score gap does not establish a reliably better model.

Compare GLM-5.3 vs GLM-5.2Full coding leaderboard

  1. 01GLM-5.3Best fit

    Z.AI · Open Weight

    61.4

    Coding score estimate

    • Highest coding estimate among models that meet every stated constraint
    • 1M context
    • Open weights, so it can run on your own hardware

    Estimate sources: Artificial Analysis, LiveBench Coding, Vals AI coding-task composite

    Model page and evidence

  2. 02GLM-5.2

    Z.AI · Open Weight

    60.8

    Coding score estimate

    • Estimate 60.8 on the same coding evidence
    • 1M context
    • Open weights, so it can run on your own hardware

    Estimate sources: LMArena Code Overall, Artificial Analysis, LiveBench Coding, Vals AI coding-task composite

    Model page and evidenceCompare with GLM-5.3

  3. 03Qwen3.8 Max

    Alibaba · Open Weight

    60.7

    Coding score estimate

    • Estimate 60.7 on the same coding evidence
    • 1M context
    • Open weights, so it can run on your own hardware

    Estimate sources: LMArena Code Overall, Artificial Analysis, LiveBench Coding, Vals AI coding-task composite

    Model page and evidenceCompare with GLM-5.3

  4. 04Hy4 preview

    Tencent · Open Weight

    59.4

    Coding score estimate

    • Estimate 59.4 on the same coding evidence
    • 1M context
    • Open weights, so it can run on your own hardware

    Estimate sources: LMArena Code Overall

    Model page and evidenceCompare with GLM-5.3

  5. 05Ornith-1.5-397B

    Ornith AI · Open Weight

    59.0

    Coding score estimate

    • Estimate 59.0 on the same coding evidence
    • 262K context
    • Open weights, so it can run on your own hardware

    Model page and evidenceCompare with GLM-5.3

Refine for your situation

Each link opens the selector with one answer changed. The address carries the answers, so your version is as shareable as this page.

A decision model maps your sentence to the selector’s questions. The shortlist comes from public evidence only, and nothing is stored.

What this shortlist rests on

The coding surface. Each model’s estimate names its own sources above; these are the public weighted benchmarks for the category.

What to verify before choosing

  • Open weights are only the first requirement. Hardware, quantization, license, and tool access need verification.
  • Composite scores are estimates. A small score gap does not establish a reliably better model.

Try three representative examples of your own work. Compare errors, time, cost, and the tools available in your actual setup.

Questions

Which is the best local LLM for coding?

On BenchLM's public evidence, GLM-5.3 by Z.AI has the highest coding score estimate among models that meet the page's default constraints (61.4). Ordered by task score under the stated constraints. Open weights are only the first requirement. Hardware, quantization, license, and tool access need verification.

What are the alternatives to GLM-5.3 for local LLM for coding?

GLM-5.2 (60.8), Qwen3.8 Max (60.7), Hy4 preview (59.4), Ornith-1.5-397B (59.0) follow on the same evidence. A small gap does not establish a reliably better model; compare them on three representative examples of your own work.

How does BenchLM pick the best local llm for coding?

The page runs the LLM Selector with fixed answers: a builder choosing by accuracy, running on their own hardware, ordinary input size. The selector uses the coding evidence surface, filters by the stated constraints, and orders by that evidence. It never adds a hidden fit score or a bonus for open weights or reasoning style.

Can I change the constraints?

Yes. Every link under "Refine" opens the selector with one answer changed, and the address carries the answers so a result can be shared or reopened against the current dataset.

Method: bench-align-v5.5-2026-09-04. Read the methodology and benchmark confidence pages for how scores and verification statuses are produced.

Watch the local LLM for coding shortlist

One weekly email when rank, price, or benchmark evidence changes make this shortlist worth revisiting.

Read a sample issue

Join 2,000+ readers.