ROBOTWAR.ioOSINT research running on the inference APIbeta test
Hetzner Inference API — Status
Independent measurement of https://inference.hetzner.com/api/v1 · one short question per model every 15 minutes · published 2026-08-23 08:18 CEST · all times on this page are Amsterdam (CET/CEST)
Size and precision are read from each model's own name; anything else needs a cited source. Weights are parameters times bytes per parameter — serving needs more than that.
Qwen3.6-35B-A3B-FP8
healthy
Qwen
size
35B total, 3B active
routing
mixture of experts · 9% active per token
served at
FP8 (1 byte per parameter)
weights
≈ 35 GB
health check
0.40 s typical
throughput
86.1 tokens/s
Quantisation, active parameters, parameters read from the model name.
Qwen3.8-27B
healthy
size
27B parameters
health check
0.68 s typical
throughput
32.9 tokens/s
Parameters read from the model name.
What is measured
Synthetic probes against the live endpoint — not this section's research workload, which is reported separately in the performance report. Two request sizes are sent, and they are never mixed in a chart or a percentile.
request
kind
measurements
what is sent
Health check
ping
653
A fixed one-line question with a per-request code in it, capped at 64 tokens. Sent to every model every 15 minutes. Identical every time, so this line is a true like-for-like comparison over time.
Working request
work
163
A short passage to summarise in about 60 words plus three key terms, capped at 400 tokens. Sent hourly. It produces roughly ten times the output of the health check, which is why it is slower and why it is charted separately.
Response time — health check
The identical short request, every measurement, at the time it was taken. Timeouts are excluded: a timeout measures the deadline, not the model.
Each point is one measured reply, plotted at the time it was taken. A break in a line is a stretch nobody measured, not a fast reply.
Response time — working request
The larger request, on the same scale idea but its own chart. It asks for about ten times the output, so its times are not comparable with the health check above.
Each point is one measured reply, plotted at the time it was taken. A break in a line is a stretch nobody measured, not a fast reply.
Speed by request size
The last column normalises the two: milliseconds per output token. When those agree, the difference in raw seconds is the size of the request, not the endpoint.
model
request
n
tokens out (median)
p50
p90
slowest
per output token
Qwen/Qwen3.6-35B-A3B-FP8
Health check
337
11
0.40 s
0.92 s
59.8 s
36 ms
Qwen/Qwen3.6-35B-A3B-FP8
Working request
84
100
2.92 s
10.9 s
48.5 s
29 ms
Qwen3.8-27B
Health check
316
11
0.68 s
2.65 s
33.0 s
62 ms
Qwen3.8-27B
Working request
79
115
6.04 s
14.2 s
112.6 s
53 ms
Uptime
Two windows on purpose. 48 hours answers 'is it healthy now'; a week answers 'can I plan on it'. A single average hides a model that broke yesterday.
Qwen/Qwen3.6-35B-A3B-FP8
48 hours94.4%1 week95.9%
Qwen3.8-27B
48 hours95.5%1 week90.0%
Uptime over time
One point per day (Amsterdam) per model. A day nobody measured leaves a gap rather than a zero.
Two or more consecutive bad cycles. A single blip is not an incident.
model
from
length
mostly
measurements
Qwen/Qwen3.6-35B-A3B-FP8
2026-08-23T05:00:02Z
30 min
unservable
3
Qwen3.8-27B
2026-08-23T04:00:04Z
29 min
unservable
3
Qwen3.8-27B
2026-08-21T08:18:02Z
12 min
unservable
2
Qwen/Qwen3.6-35B-A3B-FP8
2026-08-21T07:16:03Z
88 min
failed
7
Qwen3.8-27B
2026-08-20T20:16:03Z
15 min
failed
2
Qwen3.8-27B
2026-08-20T11:31:22Z
15 min
failed
2
Qwen3.8-27B
2026-08-19T20:32:03Z
14 min
failed
2
Qwen/Qwen3.6-35B-A3B-FP8
2026-08-19T13:01:02Z
30 min
failed
3
Qwen3.8-27B
2026-08-19T10:01:04Z
164 min
failed
12
Qwen3.8-27B
2026-08-19T06:45:03Z
15 min
failed
2
Changes to the model list
A model withdrawn from the list is not an outage — watch this before reading one.
The list of models has not changed in this window.
Were we looking?
Measurement coverage per hour.
48 of the last 48 hours were measured. Amber is a partly-measured hour and pale grey means nobody was looking — neither means anything was down.
About this endpoint
Service status
Beta. Hetzner offers this endpoint under Experiments, so it is a trial service rather than a general-availability product — expect the model list, and the numbers below, to move.
Endpoint
https://inference.hetzner.com/api/v1
Probe interval
every 15 minutes, per model
Gateway cut-off
300.0 s — a request past this is ended by the gateway
Streaming
not used, so latency below is the whole response, not time to first token
Window
30 days (2026-07-24 08:18 CEST onward)
Rate-limit headers
not returned by this endpoint — our own client-side budget is the limit