mrkeyoor.com_
Fri 02 Oct 14:58 UTC
AI Toolsevaluationupdated 02 Oct 2026

cek-probe-model review

Cek Probe Model is an Indonesian-language set of Python scripts for checking OpenAI-compatible model endpoints; it does not provide English documentation. It lists available models, sends identity or logic prompts, and prints a summary intended to flag unavailable endpoints or suspicious model labels.

Verdict

Our run installed 35 packages in 13 seconds and passed its 4-second build, but it found no test target for the code that assigns model-trust labels. Use Cek Probe Model to learn whether an authorized OpenAI-compatible endpoint responds and to collect a few clues. Do not use GENUINE / HIGH-TIER as evidence that the advertised model actually served the request.

We ran it

Lab card: what happened when we ran cek-probe-modelScreenshot of cek-probe-model (github.com/yureii1996/cek-probe-model)
Install✓ · 13s35 packages · 37 MB
Build✓ · 4s
Testsn/ano test script
Known vulns0(pip-audit)
Repo6 files~706 lines of source · 0 MB · 0 CI workflows

Answers from our run

Does cek-probe-model build from source?

Dependencies installed in 13 seconds (35 packages), and the build succeeded in 4 seconds. We cloned commit d0ae022 into a clean Debian container with 3 CPUs and no project-specific setup.

Does cek-probe-model have tests you can run?

Not through a standard command: the project exposes no test script or target that our harness could run.

Does cek-probe-model have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use cek-probe-model?

Buyers who need proof that a provider served a named model: the main script relies on the model describing its own identity, which is not independent verification.

What are the alternatives to cek-probe-model?

Promptfoo, LiteLLM, Langfuse. Our run installed 35 packages in 13 seconds and passed its 4-second build, but it found no test target for the code that assigns model-trust labels.

Setup4/513-second install and simple environment variables
Docs2/5Short Indonesian guide covers use and billing, but not validity limits
Community2/5184 stars, no open items, and a last push on September 11, 2026
Maturity1/5Six files, no license, releases, tests, CI, or structured output

Who it’s for

Indonesian-speaking developers who need a quick manual smoke check for an OpenAI-compatible endpoint.
API resellers checking whether each advertised model returns a chat-completion response.
Hobbyists willing to read the code and treat its identity labels as clues rather than proof.
Operators who can run probes only against endpoints and keys they are authorized to test.

Who it’s NOT for

Buyers who need proof that a provider served a named model: the main script relies on the model describing its own identity, which is not independent verification.
Evaluation teams comparing model quality: three familiar logic questions are not a representative or resistant benchmark.
CI pipelines that need tested machine-readable results: our scan found no tests target, no tests directory, and 0 CI workflow files, while the scripts print tables for a person.
English-only teams: the README and most code comments are Indonesian, with no English guide.
Providers without OpenAI-style /models and /chat/completions routes, since the scripts assume those response shapes.

Setup reality

Our fresh Debian sandbox installed commit d0ae022 in 13 seconds, adding 35 Python packages and using 37 MB. The build passed in 4 seconds. There was no tests script or target, so tests were skipped; pip-audit reported 0 known vulnerabilities.

The generic script needs a BASE_URL and one or more API_KEYS. Two other scripts default to InferHub and CFRouter routes and use provider-specific environment variables. Every run sends paid or quota-limited inference requests, a cost the README states directly.

The repository has 6 files, about 706 lines of source, 0 CI workflow files, no Dockerfile, and no tests directory. It has no published release or recognized license. Our sandbox did not call a provider because it had no secrets, so the clean build does not validate any model verdict.

The script checks endpoint behavior, not model identity

Cek Probe Model points at an OpenAI-compatible base URL, fetches /models, lets you select labels, and sends chat-completion requests. That is enough to answer a useful operational question: does this key reach this advertised route and return the expected response shape? The main script masks keys in console output and reads secrets from environment variables or a local .env file.

Its larger claim does not hold up. The identity probe asks the model to name its creator, architecture, cutoff date, and whether a system prompt made it pretend to be something else. A backend can answer incorrectly, follow an injected identity, refuse, or return whatever the provider instructed it to return. Comparing that text with a label such as DeepSeek or GPT produces a suspicion signal, not independent evidence of the serving model.

Three familiar puzzles cannot separate high-tier models from substitutes

The second probe asks which is larger, 9.11 or 9.9, counts the letter r in strawberry, and solves the bat-and-ball question. The script turns those replies into a score out of 3, then prints labels such as GENUINE / HIGH-TIER, SUSPECT (MEDIUM), or SUSPECT DOWNGRADE. This is a very small capability check with well-known answers.

The parser also weakens its own result. The strawberry check passes whenever the character 3 appears anywhere in the logic response. The requested answer format itself includes a 3: label for the third question, so a neatly formatted wrong reply can still receive that point. The ball check similarly looks for 5 anywhere instead of parsing the third answer. There are no unit tests covering these scoring rules.

A model that passes all three questions may still be a cheaper substitute, and a named model can miss a brittle prompt without being fake. Provider verification needs evidence outside the model's generated words, such as provider-signed metadata, controlled server logs, a much broader blinded evaluation, or a contractual audit path. This repository supplies none of those.

What happened when we ran it

Our sandbox installed commit d0ae022 in 13 seconds, adding 35 packages and consuming 37 MB. The build completed in 4 seconds. The checkout contained 6 files, about 706 lines of source, and rounded to 0 MB in the lab report. We used Python 3.12 in a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, and no secrets.

There was no tests script or target, so tests were skipped. Pip-audit reported 0 known vulnerabilities. The repository also had 0 CI workflow files, no Dockerfile, and no tests directory. Those measurements show that the code is cheap to install and did not trip the dependency audit. They do not validate the probe, because we did not have an authorized endpoint or API key to query.

The absence of secrets was appropriate for this run. The README warns that probes send inference requests and can consume provider quota or billing. It also tells users to test only endpoints and keys they own or have permission to use. That is the right boundary for a tool that can loop across every model returned by an account.

The three scripts serve different providers and produce human-readable tables

cek_model_fast.py is the general option. It accepts a base URL and several keys, lets a person choose all models or a subset, runs 2 prompts per model, and prints a final table. The key mask retains the first 8 and last 4 characters for long values. That helps distinguish credentials in a local terminal, though shared logs still reveal a stable key fingerprint.

cek_model_w_waiting.py defaults to InferHub. It sends a simple greeting, waits 3 seconds between models, and retries selected server errors up to 3 times with increasing delays. cek_model_waiting_3s.py defaults to CFRouter and asks each model to print 407. Despite its filename and comments, the current loop sleeps for 1 second between models. Neither script writes JSON or returns a failure code designed for CI.

The API assumptions are narrow and clear in code: /models must return a data list, and /chat/completions must return choices[0].message.content. A compatible proxy may satisfy that shape. A provider using a different model-list schema, responses endpoint, asynchronous job, or non-chat API will need changes.

Indonesian documentation is enough to run it, not to assess it

The README is 1,257 bytes and covers installation, environment variables, the 3 scripts, authorization, secret handling, and the possibility of quota charges. It is written in Indonesian and has no English section. The source mixes Indonesian and English names, so an English-speaking Python developer can still trace the requests, but the intended audience is clear.

What is missing is more important than another setup example. There is no explanation of false positives, no threat model for a dishonest proxy, no calibration dataset, and no statement that self-reported identity cannot authenticate a backend. The repository also has no recognized license, so reuse terms are unclear even though the source is public.

One day of code history is too little for a trust tool

GitHub showed 184 stars, 0 combined issues and pull requests, and a last push on September 11, 2026. The repository was created on September 10 and has no published release. Zero open items do not prove stability when there is almost no public issue history, no test suite, and no tagged artifact.

Keep the scripts if you want a small endpoint checklist and are willing to inspect the raw replies yourself. Rename the printed verdicts in your own workflow or ignore them. The honest output is “this route answered these prompts,” which is useful. Calling that response proof of model authenticity asks the script to establish something it never observes.

Alternatives

ProjectWhat it isPick it when
Promptfoo gh↗A configurable prompt, model, agent, and security evaluation framework with CI support.pick this instead when you need repeatable test cases, assertions, comparisons, and automated reports.
LiteLLM gh↗An AI gateway and SDK for calling many providers through common interfaces with routing and logging.pick this instead when endpoint health, fallback routing, usage tracking, and provider normalization are the job.
Langfuse gh↗A self-hostable platform for tracing and evaluating language-model applications.pick this instead when you need production traces and evaluation history rather than a terminal spot check.

What people are saying

  1. [velocity-scout] yureii1996/cek-probe-model

Sources

  1. Cek Probe Model repository
  2. Cek Probe Model README
  3. Fast model probe source
  4. InferHub waiting probe source
  5. CFRouter waiting probe source

More ai tools reviews

xialingguo-ip · reelbench-skills · SoL-Pi · flybook · crypto-rag · anything2explainer · the whole board →