The script checks endpoint behavior, not model identity
Cek Probe Model points at an OpenAI-compatible base URL, fetches /models, lets you select labels, and sends chat-completion requests. That is enough to answer a useful operational question: does this key reach this advertised route and return the expected response shape? The main script masks keys in console output and reads secrets from environment variables or a local .env file.
Its larger claim does not hold up. The identity probe asks the model to name its creator, architecture, cutoff date, and whether a system prompt made it pretend to be something else. A backend can answer incorrectly, follow an injected identity, refuse, or return whatever the provider instructed it to return. Comparing that text with a label such as DeepSeek or GPT produces a suspicion signal, not independent evidence of the serving model.
Three familiar puzzles cannot separate high-tier models from substitutes
The second probe asks which is larger, 9.11 or 9.9, counts the letter r in strawberry, and solves the bat-and-ball question. The script turns those replies into a score out of 3, then prints labels such as GENUINE / HIGH-TIER, SUSPECT (MEDIUM), or SUSPECT DOWNGRADE. This is a very small capability check with well-known answers.
The parser also weakens its own result. The strawberry check passes whenever the character 3 appears anywhere in the logic response. The requested answer format itself includes a 3: label for the third question, so a neatly formatted wrong reply can still receive that point. The ball check similarly looks for 5 anywhere instead of parsing the third answer. There are no unit tests covering these scoring rules.
A model that passes all three questions may still be a cheaper substitute, and a named model can miss a brittle prompt without being fake. Provider verification needs evidence outside the model's generated words, such as provider-signed metadata, controlled server logs, a much broader blinded evaluation, or a contractual audit path. This repository supplies none of those.
What happened when we ran it
Our sandbox installed commit d0ae022 in 13 seconds, adding 35 packages and consuming 37 MB. The build completed in 4 seconds. The checkout contained 6 files, about 706 lines of source, and rounded to 0 MB in the lab report. We used Python 3.12 in a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, and no secrets.
There was no tests script or target, so tests were skipped. Pip-audit reported 0 known vulnerabilities. The repository also had 0 CI workflow files, no Dockerfile, and no tests directory. Those measurements show that the code is cheap to install and did not trip the dependency audit. They do not validate the probe, because we did not have an authorized endpoint or API key to query.
The absence of secrets was appropriate for this run. The README warns that probes send inference requests and can consume provider quota or billing. It also tells users to test only endpoints and keys they own or have permission to use. That is the right boundary for a tool that can loop across every model returned by an account.
The three scripts serve different providers and produce human-readable tables
cek_model_fast.py is the general option. It accepts a base URL and several keys, lets a person choose all models or a subset, runs 2 prompts per model, and prints a final table. The key mask retains the first 8 and last 4 characters for long values. That helps distinguish credentials in a local terminal, though shared logs still reveal a stable key fingerprint.
cek_model_w_waiting.py defaults to InferHub. It sends a simple greeting, waits 3 seconds between models, and retries selected server errors up to 3 times with increasing delays. cek_model_waiting_3s.py defaults to CFRouter and asks each model to print 407. Despite its filename and comments, the current loop sleeps for 1 second between models. Neither script writes JSON or returns a failure code designed for CI.
The API assumptions are narrow and clear in code: /models must return a data list, and /chat/completions must return choices[0].message.content. A compatible proxy may satisfy that shape. A provider using a different model-list schema, responses endpoint, asynchronous job, or non-chat API will need changes.
Indonesian documentation is enough to run it, not to assess it
The README is 1,257 bytes and covers installation, environment variables, the 3 scripts, authorization, secret handling, and the possibility of quota charges. It is written in Indonesian and has no English section. The source mixes Indonesian and English names, so an English-speaking Python developer can still trace the requests, but the intended audience is clear.
What is missing is more important than another setup example. There is no explanation of false positives, no threat model for a dishonest proxy, no calibration dataset, and no statement that self-reported identity cannot authenticate a backend. The repository also has no recognized license, so reuse terms are unclear even though the source is public.
One day of code history is too little for a trust tool
GitHub showed 184 stars, 0 combined issues and pull requests, and a last push on September 11, 2026. The repository was created on September 10 and has no published release. Zero open items do not prove stability when there is almost no public issue history, no test suite, and no tagged artifact.
Keep the scripts if you want a small endpoint checklist and are willing to inspect the raw replies yourself. Rename the printed verdicts in your own workflow or ignore them. The honest output is “this route answered these prompts,” which is useful. Calling that response proof of model authenticity asks the script to establish something it never observes.

