LLM application code security

Your model's output is untrusted input. Find where your code forgets it.

Static analysis of the code around your LLM — the code that builds the prompt, calls the model and does something with the answer. Not the model. Not the prompt's resistance to a jailbreak. The code.

The free plan covers 3 projects and 30 analyses a month, with no limit on the number of users and no credit card. This analysis runs on it.

The problem

A model is not a library. It is an untrusted interpreter: you hand it a string and it hands one back. Whatever influenced that string — the user's own prompt, a page it was shown, a document it was given, another tool's result — influenced the answer.

The answer is harmless until the code gives it a capability. The danger is not that a model said something; it is that a program believed it. Every place a program treats generated text as code, a command, a query, a path or a URL is an injection with one extra hop — and unlike the question of whether a prompt can be jailbroken, that hop is visible in the source.

What Metivect looks for

Where model output reaches something that acts. These are the engine's own categories, severities and CWE identifiers.

Code execution — CWE-94
critical
eval, exec, compile, new Function, pickle.loads, yaml.load — the answer becomes a program.
A shell — CWE-78
critical
os.system, subprocess, child_process.exec, execSync — the answer becomes a command.
A SQL query — CWE-89
high
execute, raw, createQuery — text-to-SQL is this pattern, and DROP TABLE is a plausible completion.
A filesystem path — CWE-22
high
open, readFile, shutil.rmtree, unlink — ../../etc/passwd is a plausible completion too.
An outbound request — CWE-918
high
requests, httpx, fetch, axios, urlopen — server-side request forgery, chosen by the model.

A concrete example

Nine lines, and the whole shape of the risk.

from openai import OpenAI
client = OpenAI()

def run_plan(task: str):
    answer = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": task}],
    )
    exec(answer.choices[0].message.content)
Source
client.chat.completions.create(…) — a call that returns model output. The analyzer binds the variable it was assigned to, which is what lets it follow the value across the function without inter-procedural analysis.
Flow
answer holds attacker-influenced text. Nothing between the call and the sink inspects it, constrains it or maps it onto a fixed set of operations.
Sink
exec(…) — the text becomes a program in the application's own process, with its credentials, its network and its filesystem.
Why it matters
Anything that can influence task — directly, or through a document the model was shown — chooses what runs. This is remote code execution with an extra hop.

Run against the current engine, that snippet produces one finding: llm-output-to-code-execution, CWE-94, severity critical, confidence high, on the exec line. Confidence is high because the variable is bound to an actual completion call. A name that merely looks like a model response — llm_response, reply — is reported at medium confidence instead, because that is a guess and saying otherwise would be a lie about evidence.

What this is not

Not a test of the model
We do not probe the model, red-team it, or evaluate how it behaves under pressure. Nothing here says whether a prompt can be jailbroken: that is not decidable from the source, and claiming it would be the kind of statement this engine exists to avoid.
Not complete prompt-injection detection
Indirect injection is reported where untrusted content demonstrably reaches a prompt. A channel we cannot see in the code is a channel we do not report.
Not a guarantee about an agent
We report that an agent was granted a tool that runs code or commands. What the agent then does at run time is outside anything static analysis can decide.
Complementary to runtime AI security
Guardrails, evaluation harnesses and red-teaming answer questions this cannot, and this answers questions they cannot. Neither replaces the other.

Other LLM-specific risks in the code

Each one was triggered on a sample against the current engine before it was written here.

Secrets and personal data interpolated into a prompt — CWE-200
An api_key, a password, an IBAN or a medical record placed into the text you send. Reported where a real prompt context is being built, not on any f-string that happens to contain a key.
An agent handed a tool that runs code or commands — CWE-250
PythonREPLTool, ShellTool, a code interpreter, a terminal. Granting one of these turns every prompt-injection question into a remote-code-execution question.
A framework safeguard switched back off — CWE-502
allow_dangerous_deserialization=True, trust_remote_code=True, dangerouslyAllow… — a framework disabled it because it was unsafe.
Provider safety controls disabled — CWE-693
BLOCK_NONE, safety_settings=None, moderation=false.
A prompt or agent definition fetched at run time — CWE-829
hub.pull(…), a prompt template loaded over HTTP. Whoever controls that endpoint controls the instructions.
Untrusted content loaded straight into a prompt — CWE-77
A scraped page, a PDF, a repository, a wiki. This is the indirect injection channel, recognised by where the text came from rather than by what it says.
A model API key written into the source — CWE-798
sk-…, sk-ant-…, AIza…, hf_… committed into the repository.
A withdrawn or superseded model pinned in code — CWE-1104
text-davinci-*, gpt-4-0314, claude-instant — weaker instruction hierarchy, no system-prompt separation.

Why you should believe any of this

A published benchmark
Precision, recall and F1 per CWE, per language and per analysis layer, over a versioned, hand-labeled corpus. See the figures.
A published perimeter
Every layer we do not measure is named, with the reason. Read the methodology.
Validation on public code
Repositories and commits analysed, languages encountered, completion rate and determinism — volumes and rates, not accuracy claims. See the runs.
Every finding carries its evidence
The file, the line, the matched construct and the CWE. A finding you cannot check is a finding you should not act on.

Provider-agnostic by construction. The analyzer recognises the shape — something produces text, something consumes it — rather than one SDK's spelling. It names the framework it saw when it can (OpenAI, Anthropic, Google, Azure OpenAI, Mistral, Bedrock, Ollama, LiteLLM, LangChain, LangGraph, LlamaIndex, Semantic Kernel, AutoGen, CrewAI, Hugging Face) so a finding can say where it came from. That is a label on the finding, not a per-SDK integration, and a framework absent from that list is analysed exactly the same way.

See Metivect reason about a real-world system

Watch the interactive demo, then request a guided trial for your team.