A claim that one model produces safer code is only useful when you can inspect how it was tested. Prompt wording, model version, tools, repository context and evaluation criteria all influence the result.
This is an evaluation method, not a report of a completed experiment. No comparative detection rate or model ranking is asserted here.
Define the decision before the dataset
Choose a question your team can act on. Are you selecting a coding assistant, testing a model update, or deciding how much security review a generated feature needs? A result on small standalone functions does not automatically answer a question about a multi-tenant SaaS repository.
List the stack, languages and security boundaries that matter to your product. Include tasks that touch identity, access to customer records, payment state and external network requests. Keep evaluation fixtures synthetic and free of production credentials or customer data.
Make the experiment reproducible
Record the exact model identifier, date, generation settings, system instructions, tools and input repository revision. Save the original prompt and generated patch. Repeat the tasks enough to observe variation; one lucky output cannot establish a reliable success rate.
Use the same task and acceptance criteria for each candidate. If one candidate receives additional context, tools or retries, report that difference instead of presenting the result as a direct model-only comparison.
Separate functionality from security
A patch can pass its functional tests while violating a permission boundary. For each task, define a successful allowed action and a rejected unauthorized action. For example, a record owner should be able to update their record while a different account should fail through the same endpoint.
The Next.js data security guide provides relevant framework context for server-side entry points and client boundaries. Adapt tests to the version and architecture you actually use.
Use automated findings as evidence to review
Run static checks and dependency audits against each generated artifact, then have a reviewer confirm the important findings. A warning count mixes true positives, duplicates, advisory issues and false positives. It is not a security ranking by itself.
Record missed issues as well as detected issues. When a scanner reports nothing, test the boundary anyway. A clean result means the configured check did not report a match; it does not establish that the output is secure.
Report outcomes with their limitations
Publish the task definitions, scoring rubric, versions and reproducible fixtures alongside aggregate results. Separate functional failures, confirmed security failures and uncertain cases. Explain whether reviewers knew which model produced each patch and how disagreements were resolved.
Avoid extrapolating a small internal evaluation to all languages, prompts or production systems. Rerun relevant cases after a model, dependency or framework change. If a result cannot be reproduced, label it as an observation rather than a benchmark.
Turn the evaluation into a release control
Keep the strongest denied-action tests in your application test suite. Use the AI code security review workflow to connect generated patches to source review and release checks. For dependencies, review the npm auditing documentation and the resolved versions in your lockfile.
Bryxe can contribute supported code and dependency findings to this process. It does not establish a universal winner between models, and a scan cannot replace independent verification of the behavior your application requires.
