LOADTEST — Insika
How to load-test the engine, how to compare it apples-to-apples against another gateway, and how to read the numbers — covering the SQLite-topology question (does one box hold up?) and loadtest parity. For deploy/env details see DEPLOY.md.
The whole point: the engine exposes POST /v1/responses as an SSE drop-in of
the OpenClaw gateway. Same contract → the same load tools work against either side,
so you can measure the engine you are about to ship against the gateway it replaces.
There are four scripts, each answering a different question:
| Script | Question it answers | Needs a provider? |
|---|---|---|
scripts/bench_store.rb |
Does SQLite (WAL) hold up N processes writing the same file? | No |
scripts/loadtest.rb |
End-to-end: TTFB/total/tokens/cache/error against /v1/responses |
Yes |
scripts/loadtest-local.sh |
Single-proc baseline vs N-worker multi-proc on one box | Yes |
scripts/loadtest_session.rb |
A full multi-message session (CEP, searches, FAQ) under C concurrent sessions — direct to the engine (--surface engine, stream vs steer) or through the consumer’s real ingress (--surface web, the consumer’s widget API) |
Yes |
insika soak |
Does the deploy degrade over 72 h of steady load? | Yes |
The first four take --help / -h; the soak is a shipped command (insika soak --help) rather
than a repo script. Bursts and uptime are different questions: a wave driver
measures a burst, and the soak’s arrival process measures degradation over
uptime — the load-test table above deliberately stops where
Soak begins.
1. bench_store.rb — SQLite write ceiling (no provider)
Isolates “does SQLite survive multi-process?” from LLM-provider noise. N processes
hammer writes against the same file using the engine’s real production config
(WAL + busy_timeout + BEGIN IMMEDIATE + in-process semaphore). It uses a fresh
temp db per round, so it never touches your INSIKA_DB.
bundle exec ruby scripts/bench_store.rb [PROCS_CSV] [WRITES_PER_PROC]
# defaults: 1,2,4,8 procs, 2000 writes each
bundle exec ruby scripts/bench_store.rb 1,2,4,8 3000
Output columns: procs | wall(s) | writes/s | p50(ms) | p95(ms) | max(ms) | locked.
What to look for: aggregate writes/s stays roughly flat as procs grow (the
WAL “1 writer at a time” ceiling), and locked (i.e. database is locked) is 0
— the busy_timeout turns contention into tail latency (max), not errors. A real
turn is provider-bound (seconds) and does only a handful of writes, so the workload
sits ~100× below this ceiling. See DEPLOY.md for the measured numbers.
2. loadtest.rb — end-to-end against /v1/responses (with provider)
Hits POST /v1/responses (SSE) directly — the production path
(a consumer app, e.g. WhatsApp, → engine). Standard library only. Fires agents × concurrency ×
iterations turns in waves of concurrency, and per turn records TTFB (time to
first SSE byte), total time, and the usage block (tokens + cache hit) of the last
frame that carries it.
INSIKA_URL=http://localhost:9292 \
INSIKA_GATEWAY_TOKEN=xxx \
bundle exec ruby scripts/loadtest.rb \
--agents demo,my-store --concurrency 16 --iterations 3 \
--message "hi, how are you?"
Runs against a local server or a remote one (e.g. Railway) — just point
INSIKA_URL at it.
Flags
| Flag | Default | Meaning |
|---|---|---|
--agents a,b,c |
demo |
comma-separated agent ids (mapped to model: insika:<agent>) |
--concurrency N |
8 |
concurrent turns per wave |
--iterations N |
1 |
number of waves per agent |
--message TEXT |
greeting | user message sent every turn |
--timeout SECONDS |
120 |
per-request read timeout |
--ports 9292,9293 |
— | round-robin across local processes (see §4) |
--same-user |
off | boolean toggle — reuse the same user per agent to measure the hot-conversation cache (legacy --same-user 1 / --same-user 0 still work) |
--dry-run |
— | print the plan + one sample request (masked token) and exit; sends no traffic |
--help |
— | show usage and exit |
Environment
| Env | Default | Meaning |
|---|---|---|
INSIKA_URL |
http://localhost:9292 |
base URL of the engine |
INSIKA_GATEWAY_TOKEN |
falls back to ADMIN_TOKEN, then local-demo |
Bearer for /v1/responses |
DEEPSEEK_API_KEY |
— | must be configured on the server for real turns (not read by the client) |
Use --dry-run to sanity-check your flags/URL/token before firing real traffic
(and to confirm the request body without needing a running server):
INSIKA_URL=http://localhost:9292 INSIKA_GATEWAY_TOKEN=xxx \
bundle exec ruby scripts/loadtest.rb --agents demo --concurrency 16 --dry-run
--same-user and the cache
By default each turn uses a distinct user (loadtest-<agent>-<idx>), so every
turn is a cold conversation. With --same-user all turns for an agent share one
user, exercising the warm-conversation path — watch the mean cache hit rise
and TTFB drop. Run both to bracket cold vs hot behaviour.
3. loadtest-local.sh — baseline vs multi-worker on one box
Boots Falcon with --count 1 (single-process baseline), runs the sweep, then boots
--count N (multi-process) over the same SQLite file (WAL), runs the sweep
again, and counts database is locked in each Falcon log. Since the provider is
identical across both runs, if multi-proc recovers throughput the ceiling was
CPU/event-loop (not the provider), and a locked count of 0 proves the WAL +
busy_timeout config absorbs cross-process write contention.
DEEPSEEK_API_KEY=sk-... ./scripts/loadtest-local.sh [WORKERS] [CONCURRENCY]
# defaults: 4 workers, 16 concurrency
./scripts/loadtest-local.sh 4 24
| Env | Default | Meaning |
|---|---|---|
DEEPSEEK_API_KEY |
— (required) | real turns hit the provider; also auto-sourced from .env.local |
INSIKA_GATEWAY_TOKEN |
falls back to ADMIN_TOKEN, then local-demo |
Bearer for the sweep |
PORT |
9299 |
bind port for the local Falcon |
AGENT |
demo |
agent id to load |
The final block prints the lock counts; the expected reading is 0 for both:
=== 'database is locked' (expected 0 — WAL + busy_timeout) ===
baseline (1): 0
multi (4): 0
4. Apples-to-apples: Insika vs the OpenClaw gateway
Two ways to compare, both valid because the SSE contract is identical.
4a. Ruby native (loadtest.rb) against both
Run the same loadtest.rb invocation twice — once with INSIKA_URL pointing at
the engine, once at the gateway (its /v1/responses speaks the same protocol).
Keep --agents, --concurrency, --iterations and --message identical, and use
matching agents on both sides. Compare the printed TTFB/total/cache/error lines.
4b. Comparing against an OpenClaw gateway
The engine speaks its own wire names (model: insika:<agent>, header
X-Insika-Agent), so OpenClaw’s loadtest-gateway.mjs cannot be pointed at it
unmodified. For a shadow comparison, run loadtest.rb (section 4a) against each
side with identical --agents, --concurrency, --iterations and --message,
and diff the reports. What you still need in hand:
- A bearer token accepted by each side —
INSIKA_GATEWAY_TOKENfor the engine (see DEPLOY.md), the gateway’s own token for the gateway run. - The same agent id provisioned on both sides (e.g.
demo). On the engine, provision viascripts/import_pack.rb. - The same provider (or an equivalent-latency one) behind each, otherwise you are comparing providers, not engines.
- Both endpoints reachable from where you run the client, warmed up (hit
/upon the engine first), and ideally driven from the same machine to remove network skew.
Keep every knob identical between the two runs — the only variable should be which
engine is behind /v1/responses.
5. Reading the metrics
- TTFB (time to first SSE byte) — how fast the user starts seeing a response. Dominated by provider latency + the engine’s per-turn setup (context build, policy, first model call). This is the number that most shapes perceived latency.
- total — full turn wall time including the whole tool-loop and streamed
output.
total − TTFBis roughly the streaming/tool-loop tail. - P50 vs P95 — P50 is the typical turn; P95 is the tail you actually feel
under load. A P50 that stays flat while P95 balloons as concurrency rises means
you are queueing (CPU/event-loop or writer contention) — that is the signal to
add workers (§3) or check
bench_store.rb. - mean tokens — average
total_tokens/output_tokensper turn; sanity-checks that turns did real work and lets you compare cost between runs/engines. - mean cache hit — average cached prompt tokens; should rise sharply with
--same-user 1. A high hit rate is why warm conversations are cheaper and faster. - error rate (
turns ok: X/Y (errors: N)) — non-2xx, timeouts, or connection failures. Anything above ~0 under moderate load is a red flag; inspect server logs. Noteloadtest.rberrors are transport/HTTP-level;database is lockedspecifically is counted byloadtest-local.shfrom the Falcon logs, not here. - throughput (
turns/s) — completed turns per wall second; the headline capacity number for a given concurrency.
6. Checklist — what to measure before choosing a topology
Work top-down and measure before assuming — avoid premature topology optimization.
| # | Measure | Tool | Decision it informs |
|---|---|---|---|
| 1 | SQLite write ceiling & locked count under N procs |
bench_store.rb |
Is SQLite a bottleneck at all on one box? (Expected: no.) |
| 2 | Single-proc baseline TTFB/total/P95/throughput | loadtest.rb (or loadtest-local.sh count 1) |
The reference point for everything else. |
| 3 | Multi-proc on one box: does throughput scale, locked = 0? |
loadtest-local.sh |
Do more Falcon workers help, and does the shared WAL hold? Sets WEB_CONCURRENCY. |
| 4 | Insika vs gateway, identical knobs | §4 (either method) | Is the engine at parity with the engine it replaces before cut-over? |
| 5 | Cold vs hot conversation (cache) | loadtest.rb with/without --same-user 1 |
Expected steady-state cost/latency once conversations warm up. |
| 6 | Remote (Railway) vs local | loadtest.rb with INSIKA_URL remote |
Network/deploy overhead of the real environment. |
| 7 | Degradation over uptime (72 h) | insika soak (see Soak) |
The cut argument: does latency or memory grow with uptime? Run it last — it needs the topology settled first. |
Reaching for horizontal scale is only justified after 1–3 show the single box is the limit. If it is, the paths are: sharding-by-tenant + sticky routing (recommended), LiteFS, or an optional Postgres adapter — plus Litestream for backup/DR regardless of topology. Do not skip straight to Postgres: the numbers usually show SQLite on one big box is not the bottleneck.