AI-Assisted AppSec: Turning Security Expertise Into Repeatable Workflows

GitHub Security Lab used structured AI workflows to identify Android vulnerabilities. The more important lesson is how security expertise, AI reasoning, deterministic tools and human verification can work together.

Read in: English · తెలుగు · हिन्दी

Diagram showing source code passing through an AI-assisted security investigation and evidence validation before becoming a confirmed vulnerability.

GitHub Security Lab’s Android research shows how structured taskflows can help AI investigate vulnerabilities, and why candidate findings still need technical evidence and human judgement.

Traditional security scanners are good at finding patterns they have been designed to recognise.

A security researcher does something harder. They identify exposed entry points, follow untrusted data through the application, question trust boundaries, and build a theory about how several pieces of code might combine into a real vulnerability.

GitHub Security Lab is experimenting with whether part of that investigation process can be turned into a reusable AI-assisted workflow.

On September 28, 2026, the team described using its open-source Taskflow Agent to find and report 24 vulnerabilities in Android applications. The number is noteworthy, but the more important development is how the investigation was structured.

The researchers did not simply give an LLM a repository and ask it to find bugs.

They gave the model a security investigation process to follow.

From source code to a vulnerability hypothesis

GitHub’s system uses taskflows.

A taskflow is a reusable sequence of instructions and tools that tells the AI agent how to work through a particular security task step by step.

For mobile applications, the workflow first gathers information about the application’s exposed entry points. GitHub’s current taskflows can identify Android components, deep links, URL schemes, application extensions and WebView JavaScript bridges, along with information such as permissions and export status.

The workflow can then apply mobile-specific security guidance when investigating those entry points.

From source code to a validated finding Seven-step sequence from source code through attack-surface discovery, entry-point identification, Android-specific investigation and a vulnerability hypothesis to technical testing and researcher validation. {“creator”:”TechiesJournal”,”author”:”Prasad Kukkala”,”asset”:”ai-assisted-appsec-figure-1”,”source_revision”:”ai-assisted-appsec-v1-2026-09-30”,”created”:”2026-09-30”,”rights”:”Copyright 2026 TechiesJournal. All rights reserved.”} 1 Source code Where the audit begins 2 Discover the attack surface Map what is externally reachable 3 Identify security entry points Exported activities, intents, WebView bridges 4 Apply Android-specific knowledge Mobile security patterns and mitigations 5 Develop a vulnerability hypothesis A candidate explanation, not yet proof 6 Test whether the behaviour is real Technical validation of the hypothesis 7 Researcher validates the finding Human judgement confirms the result The model provides steps 2 to 5. Steps 6 and 7 still require technical and human evidence. TECHIESJOURNAL
Figure 1: A structured AI-assisted audit separates attack-surface discovery, investigation, hypothesis generation and technical validation rather than treating the model’s output as a confirmed finding.
Accessible text alternative for Figure 1

Seven-step sequence: 1. Source code, where the audit begins. 2. Discover the attack surface. 3. Identify security entry points such as exported activities, intents and WebView bridges. 4. Apply Android-specific knowledge. 5. Develop a vulnerability hypothesis, a candidate explanation and not yet proof. 6. Test whether the behaviour is real. 7. Researcher validates the finding. A closing note states the model provides steps 2 to 5, while steps 6 and 7 still require technical and human evidence.

The LLM is important, but it is only one part of this system.

The other important part is the investigation method wrapped around it.

Security expertise has moved partly into the workflow

Traditional security tools also contain expert knowledge.

A static-analysis rule might look for attacker-controlled input reaching a dangerous function. A CodeQL query can model data moving from a source to a sensitive sink. A secret scanner knows the forms that particular credentials may take.

AI-assisted auditing allows some security knowledge to be expressed differently.

Instead of encoding every idea as a deterministic query, a researcher can tell the workflow to identify externally reachable components, understand what input an attacker can control, investigate relevant vulnerability classes, follow relationships across the code, and determine whether apparently separate behaviours can be combined.

The AI model provides flexible reasoning between those steps.

This leads to an important distinction:

The capability does not come only from a better model. It also comes from giving the model a better investigation process.

That is why these taskflows are more interesting than simply asking an LLM to review source code.

Security expertise has not disappeared. Some of it has been packaged into the workflow the model follows.

A real example: following a trust boundary in OsmAnd

One of GitHub’s examples involved the Android navigation application OsmAnd.

The researchers found an exported Android activity that accepted intent data expected to come from a trusted internal path.

Because the component was externally accessible, another application could provide those values.

The investigation then followed what happened to that data.

GitHub reported that an attacker could manipulate imported settings, change network configuration and redirect map-tile or routing requests to attacker-controlled infrastructure. This could expose location information.

Following the OsmAnd trust boundary Six-step sequence from an external application through an exported Android activity, attacker-controlled intent data, changed application settings and redirected network requests to exposed location information. {“creator”:”TechiesJournal”,”author”:”Prasad Kukkala”,”asset”:”ai-assisted-appsec-figure-2”,”source_revision”:”ai-assisted-appsec-v1-2026-09-30”,”created”:”2026-09-30”,”rights”:”Copyright 2026 TechiesJournal. All rights reserved.”} 1 External application Any other app already on the phone 2 Exported Android activity A component reachable without permission 3 Attacker-controlled intent data Values meant to stay internal 4 Application settings changed Imported without the expected check 5 Network requests redirected Map-tile and routing traffic moved 6 Location information exposed To attacker-controlled infrastructure The risk comes from the chain, not from any single suspicious line of code. TECHIESJOURNAL
Figure 2: The reported OsmAnd issue depended on following attacker-controlled data across several application behaviours, not merely identifying one suspicious API.
Accessible text alternative for Figure 2

Six-step chain: 1. External application, any other app already on the phone. 2. Exported Android activity, a component reachable without permission. 3. Attacker-controlled intent data, values meant to stay internal. 4. Application settings changed, imported without the expected check. 5. Network requests redirected, map-tile and routing traffic moved. 6. Location information exposed to attacker-controlled infrastructure. A closing note states the risk comes from the chain, not from any single suspicious line of code.

This is not simply a matter of spotting one dangerous function.

The security problem emerges from understanding who can call a component, what information they can control, how that information changes application state, and what happens later because of that change.

That kind of relationship is one reason AI-assisted code investigation is attracting attention.

It does not prove that an AI system will find every vulnerability of this kind. It also does not prove that existing security techniques could not find the same issue.

It shows where flexible code reasoning may add another useful layer to an AppSec workflow.

A candidate finding is not evidence

This is also where the limitations become important.

GitHub reports that its models can misunderstand whether a suspected vulnerability is actually exploitable. They can overlook mitigating behaviour, return low-impact issues, and estimate severity incorrectly.

A code path can look dangerous while still being harmless at runtime.

For example, a model may see that attacker-controlled data can be written somewhere and conclude that another part of the application will consume it. A later runtime behaviour may give trusted data precedence instead.

The model’s explanation may sound convincing.

The vulnerability may still not exist.

For security teams, this creates a necessary boundary:

AI can generate a vulnerability hypothesis. Exploitability still needs evidence.

That evidence may come from a proof of concept, runtime testing, instrumentation, a debugger, fuzzing, or manual reproduction.

The more serious the claimed vulnerability, the more important this distinction becomes.

GitHub’s earlier results show the size of that gap

GitHub Security Lab published another useful experiment in March 2026.

Researchers ran related auditing taskflows against more than 40 repositories, primarily multi-user web applications.

The LLM initially suggested 1,003 possible issues.

After a further audit stage, deduplication and filtering, researchers manually inspected 91 candidates.

They rejected:

  • 20 as false positives.
  • 52 as too low-impact to report.

They retained 19 findings that they considered serious enough to report.

From AI suggestion to reportable vulnerability Funnel from 1,003 AI-generated candidate issues to 91 manually investigated, with 20 false positives and 52 low-impact findings removed, ending at 19 reported vulnerabilities. {“creator”:”TechiesJournal”,”author”:”Prasad Kukkala”,”asset”:”ai-assisted-appsec-figure-3”,”source_revision”:”ai-assisted-appsec-v1-2026-09-30”,”created”:”2026-09-30”,”rights”:”Copyright 2026 TechiesJournal. All rights reserved.”} 1,003 Potential issues suggested by the AI audit 91 Manually investigated by researchers 20 false positives 52 too low-impact 19 Reported vulnerabilities Raw AI output is the wrong measure of success. What matters is how many candidates survive technical verification and are worth fixing. TECHIESJOURNAL
Figure 3: GitHub’s earlier audit experiment shows the large gap between initial AI-generated candidates and vulnerabilities considered significant enough to report.
Accessible text alternative for Figure 3

Funnel from 1,003 potential AI-generated issues to 91 manually investigated candidates. A branch shows 20 false positives and 52 low-impact findings removed. The funnel ends at 19 reported vulnerabilities. A closing note states raw AI output is the wrong measure of success, and what matters is how many candidates survive technical verification and are worth fixing.

These numbers are useful because they show why raw AI output is the wrong measure of success.

An AppSec team does not benefit simply because an agent produces hundreds of findings.

The useful question is: how many survive serious technical verification and matter enough to fix?

Where does this fit beside existing AppSec tools?

AI-assisted auditing should not be treated as a replacement for every existing security technique.

Different tools provide different kinds of evidence.

Where different AppSec capabilities provide evidence
CapabilityWhat it is good at
Static analysisRepeatable checks for known code patterns and data flows
FuzzingExercising real software with generated inputs and observing runtime behaviour
AI-assisted investigationExploring code relationships and developing vulnerability hypotheses
Security researcherValidating exploitability, severity, context and business impact

These capabilities can work together.

GitHub’s September 2026 fuzzing work illustrates this approach. Its taskflow uses an LLM agent to identify promising targets, understand build systems, create fuzzing harnesses, review coverage and triage crashes, while AFL++ remains the execution engine performing the fuzzing itself.

That is a more realistic direction than expecting an LLM to replace an AppSec toolchain.

AI can help decide where and how to investigate.

Deterministic tools can test what they are designed to test.

Humans remain responsible for the security judgement built on that evidence.

The taskflow itself becomes a security engineering asset

There is another consequence if organisations begin using workflows like these regularly.

The taskflow itself becomes something that needs engineering discipline.

Changing the model, prompt, available tools or investigation sequence can change the results.

A mature implementation would therefore need to think about at least five things.

Versioning

Teams should know which model and workflow version produced a finding.

Regression testing

Known test cases can help answer whether a changed workflow still detects problems it previously found.

Model-change testing

Moving to a newer model should not automatically be treated as an improvement. Security behaviour needs to be tested again.

Permissions

An agent that only reads source code presents a different risk from one allowed to compile software, execute commands or run proofs of concept.

Evidence

A security team should be able to trace how an important finding was produced and what evidence confirmed it.

This becomes particularly important when the agent can execute tools.

GitHub recommends running its auditing taskflows in a sandboxed environment, and the Taskflow Agent documentation warns that its Docker image should not itself be considered a security boundary.

Once an AI security agent can run code, the security design must include the agent itself.

Where AI can help, and where evidence is still needed

Where AI helps in AppSec, and what still needs verification
StageAI can helpWhat still needs verification
Repository reconnaissanceYesImportant omissions
Code-path investigationYesAssumptions about actual behaviour
Vulnerability hypothesisYesReproduction
Proof-of-concept creationYesSafe execution and result validation
ExploitabilitySupport onlyRuntime evidence
SeveritySupport onlySecurity judgement
Business impactLimited contextHuman decision

This division is more useful than asking whether AI is simply “good at security.”

Different stages require different levels of confidence.

What should AppSec teams do now?

There is no need to replace existing security tools or launch an autonomous vulnerability-hunting programme simply because these experiments are producing interesting results.

A better approach is to start narrowly.

Start with one investigation task

Use AI assistance for something measurable: entry-point discovery, analysis of existing scanner findings, authorization-path investigation, fuzzing-harness creation, or another well-defined task.

Measure the complete result

Do not count raw AI findings as success.

Measure how many findings can be reproduced, how many prove meaningful, how much researcher work is saved, and which important problems the workflow misses.

Keep an evidence gate

An AI-generated finding should not become a confirmed vulnerability simply because its explanation sounds convincing.

Exploitability, severity and impact still need appropriate technical evidence.

The workflow matters more than the headline

GitHub’s 24 Android vulnerabilities are useful evidence that AI-assisted security investigation is becoming more capable.

But the more important development is the workflow behind them.

Security researchers are beginning to package parts of their investigation methods into reusable processes around AI models.

The model provides flexible reasoning. The taskflow provides structure and security knowledge. Existing security tools provide deterministic and runtime evidence. Human researchers remain responsible for deciding whether the resulting vulnerability story is actually true.

That suggests a more practical direction for AI in AppSec:

Turn good security investigation methods into repeatable workflows, use AI where flexible reasoning helps, and require evidence before trusting the result.

That is more demanding than asking an AI model to find bugs.

It is also much closer to how security engineering actually works.

References and further reading

Report a correction

Corrections go to the editor and are never published automatically. No account needed.