v0.1.3 · macOS build available now

Routing on measurement, not vibes.

Soriku benchmarks every model you connect, local or remote, and builds a capability map: measured scores per category, on your hardware, with the uncertainty of each measurement. Every task routes to the proven best fit, and you always see why.

Download · free for solo use
Runs · fully offline
Models · Ollama · Claude · OpenAI · Mistral · Groq · own
capability map
example data
model
Code gen
Reasoning
Summarise
Translate
Code review
Security
qwen2.5-coder local
87 ±3 87 ±3 14 tests · measured on your hardware
62 ±6 62 ±6 9 tests · measured on your hardware
71 ±5 71 ±5 11 tests · measured on your hardware
44 ±9 44 ±9 6 tests · measured on your hardware
83 ±4 83 ±4 12 tests · measured on your hardware
58 ±7 58 ±7 8 tests · measured on your hardware Outdated. This measurement predates the last model change.
deepseek-r1 local
74 ±5 74 ±5 10 tests · measured on your hardware
89 ±3 89 ±3 15 tests · measured on your hardware
80 ±4 80 ±4 12 tests · measured on your hardware
63 ±7 63 ±7 8 tests · measured on your hardware
76 ±5 76 ±5 10 tests · measured on your hardware
81 ±4 81 ±4 11 tests · measured on your hardware
gemma3 local
51 ±8 51 ±8 7 tests · measured on your hardware
58 ±7 58 ±7 9 tests · measured on your hardware
77 ±5 77 ±5 11 tests · measured on your hardware
84 ±4 84 ±4 13 tests · measured on your hardware
49 ±9 49 ±9 6 tests · measured on your hardware
not measured
qwen3 local
79 ±4 79 ±4 12 tests · measured on your hardware
82 ±4 82 ±4 12 tests · measured on your hardware
75 ±5 75 ±5 10 tests · measured on your hardware
70 ±6 70 ±6 9 tests · measured on your hardware
72 ±6 72 ±6 9 tests · measured on your hardware
66 ±7 66 ±7 8 tests · measured on your hardware
Claude Sonnet remote
82 ±4 82 ±4 11 tests · measured on your hardware
86 ±3 86 ±3 13 tests · measured on your hardware
not measured
88 ±3 88 ±3 14 tests · measured on your hardware
79 ±5 79 ±5 10 tests · measured on your hardware
74 ±6 74 ±6 9 tests · measured on your hardware Outdated. This measurement predates the last model change.
why this one code gen · qwen2.5-coder 87 ±3 14 tests · measured on your hardware
measured stale not measured 6 of 16 categories shown · example scorecard, not a real export
The map in action

Every task routed on measurement.

Watch the map decide. A task comes in, gets classified, and goes to the model with the highest measured score in that category. The score and its uncertainty travel with the answer.

Live routing
Step 1/5
Task
Refactor this Python function
classifier
Classifying…
qwen2.5-coder
local·7B
deepseek-r1
local·7B
gemma3
local·4B
qwen3
local·8B
Claude Sonnet
remote··
Mistral Large
remote··
Answering · qwen2.5-coder
Whycode 87 ±3measured on your hardware · 14 tests
The problem

You have local models. The intelligence between them is missing.

Ollama runs the models, LM Studio gives you a UI, LangChain chains them together. None of those tools knows which model is good at what on your hardware, in your language, for the task in front of you. So you switch manually, or you give up and pay OpenAI for everything. Soriku closes that gap.

How it works

How Soriku gets to a routing decision.

Connect your models

Local through Ollama, or remote via your own keys for Claude, OpenAI, Mistral, Groq. Anything that speaks the OpenAI API works.

Soriku learns them

A capability benchmark runs once per model, per category. Code, reasoning, translation, summarisation. The result is a scorecard you can read.

Every task goes to the best fit

Routing happens on measured data, not vibes. Simple question to a fast local model, hard reasoning to your strongest. You always see who answered, and why.

Read the full breakdown
Why it's different
01 / 03

Local-first by design

Runs entirely on your hardware. No cloud needed, no data leaves your machine. Offline is the default, not a fallback.

02 / 03

Model-agnostic

Hook up Ollama, Claude, OpenAI, Mistral, Groq, or anything OpenAI-compatible. Soriku doesn't care where a model lives, it picks on capability and cost.

03 / 03

EU-compliant by default

Built in the Netherlands, designed inside the EU AI Act. No vendor lock-in, no US CLOUD Act exposure.

In the product

What the orchestrator does.

Smart routing

The map in action: every task goes to the measured best fit, and you see why.

Multi-model answers

When the map's confidence is low, verify and ensemble modes merge the results into one answer.

Agent personas

Persistent identities, each routed to the map's strongest model for their domain.

Capability benchmark

Measure your models on your hardware, not someone's marketing slide.

MCP integration

The map in Cursor, VS Code, the terminal, anywhere MCP lives.

Team workspaces

One shared map for the whole team, one login via Simezu.

All features
Agents & learning

Agents learn your work, the router learns your models.

Learning happens in two places. Agent personas build up memory and context the more you use them. The capability map keeps re-measuring your models, so routing gets sharper over time instead of sitting on a launch-day snapshot.

Agent personas

Give any job a persistent identity: a name, a specialisation, a tone, and its own memory. Start from the library (Content Strategist, Creative Writer, AI/ML Engineer, DevOps, Data Analyst) or build your own. Each persona remembers what it has worked on and gets more useful with every interaction.

  • Persistent memory per persona, it remembers past work, not just the last message
  • Specialised routing: a coding persona leans on your strongest code models, a writer on your best prose
  • Share a persona with your team and everyone inherits the same setup
  • An interaction count on every card, so you see which agents get used and which sit idle
soriku.ai · agents
Agent personas
soriku.ai · models
Capability map

Continuous learning

Soriku never stops measuring. Every model you add is benchmarked across 16 categories on your hardware, then re-benchmarked as it updates. Routing outcomes feed back in, so the map reflects how your models really perform today.

  • 16-category capability map, re-runnable any time
  • Scores update as your models and hardware change
  • Routing improves as the map sharpens, without manual tuning
  • Every answer carries provenance: which model, which score, why
Part of Atypisch

Built alongside the rest of the stack.

Soriku is one piece of a wider ecosystem of local-first tools, all built by the same team. Single sign-on, shared billing, one philosophy.

S

Simezu

Sign-in and billing for the ecosystem

W

Wemazu

GitHub orchestration

R

Rozuro

Invoicing and CRM

S

Sumezi

Newsletters

A

Akyru

Local LLM

Stop guessing which model to use.
Route on what you measured.

Download Soriku and start routing today. Join the mailing list for release notes and the occasional deep-dive: one email when it matters.