{
"id": "chatcmpl-mock-821bae6b",
"object": "chat.completion",
"created": 1735689600,
"model": "mock-gpt-4o-mini",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello, mock model!"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 5,
"completion_tokens": 5,
"total_tokens": 10
}
}(문서는 영어)
What it does
The Mock LLM Responder generates deterministic mock LLM API responses for testing clients. Pick a scenario - echo, canned answer, streamed poem-ish lorem, HTTP 429 rate limit, HTTP 500 server error, or slow chunks - and it produces the OpenAI-style chat completion JSON, the SSE event stream (data: lines with a [DONE] terminator and chunk timing markers), and a curl command that replays the request against a local stub. No API key, no account, no network: everything is computed in your browser from a fixed seed, so the same configuration always produces byte-identical output - exactly what a reproducible client test suite needs.
How to use it
- Choose a Scenario from the dropdown (e.g.
Streamed loremfor a chunked stream,Error 429to test backoff). - Set the Model name and Max tokens cap your client will send.
- Type the User prompt - the
echoscenario returns it verbatim as the assistant message. - Switch the Output tabs between Completion JSON, SSE stream, and curl, and copy whichever you need.
- Use the timing readout (
first byte/inter-chunk) to make your stub sleep like a real model, or Copy share link to send the exact configuration to a teammate.
Examples
Echo scenario (default settings) - completion JSON:
{
"id": "chatcmpl-mock-821bae6b",
"object": "chat.completion",
"created": 1735689600,
"model": "mock-gpt-4o-mini",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hello, mock model!" },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 5, "completion_tokens": 5, "total_tokens": 10 }
}
Streamed lorem - SSE stream (abridged). The leading : comment is the chunk timing marker:
: mock scenario=streamed-lorem first-byte=60ms inter-chunk=40ms
data: {"id":"chatcmpl-mock-36156cc1","object":"chat.completion.chunk","created":1735689600,"model":"mock-gpt-4o-mini","choices":[{"index":0,"delta":{"role":"assistant","content":"comet nebula prism comet "},"finish_reason":null}]}
data: {"id":"chatcmpl-mock-36156cc1","object":"chat.completion.chunk","created":1735689600,"model":"mock-gpt-4o-mini","choices":[{"index":0,"delta":{"content":"vector prism lumen\nvector "},"finish_reason":null}]}
...
data: {"id":"chatcmpl-mock-36156cc1","object":"chat.completion.chunk","created":1735689600,"model":"mock-gpt-4o-mini","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":5,"completion_tokens":63,"total_tokens":68}}
data: [DONE]
Error 429 - completion JSON:
{
"error": {
"message": "Rate limit reached for the mock model. Please retry after 1 second.",
"type": "rate_limit_error",
"code": "rate_limit_exceeded"
}
}
curl tab - replay the request against a local stub on port 8080:
# Local stub: reply 200 with the body shown in the JSON tab.
curl -s http://localhost:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"mock-gpt-4o-mini","messages":[{"role":"user","content":"Hello, mock model!"}],"max_tokens":64}'
Good to know
- Deterministic by design: ids, lorem text, and timestamps come from a seed derived from your scenario, model, and token cap - never from the clock or
Math.random(). Same settings in, same bytes out. Thecreatedfield is the fixed epoch1735689600(2025-01-01). - Chunks reassemble exactly: every SSE delta’s content, concatenated, equals the completion’s message content - so you can assert stream parsing against the JSON tab.
- Token math: ~4 characters per token, floor of 1; hitting the
Max tokenscap flipsfinish_reasonto"length". - Private: runs 100% client-side - no request leaves your browser.
- Related tools: LLM Cost Calculator, Token Estimator, HTTP Methods Tester.