GitHub Security Lab యొక్క Android పరిశోధన, structured taskflows AI కి vulnerabilities investigate చేయడంలో ఎలా సహాయపడతాయో, మరియు candidate findings కు ఇంకా ఎందుకు technical evidence మరియు human judgement అవసరమో చూపిస్తుంది.
Traditional security scanners తాము గుర్తించడానికి design చేయబడిన patterns ను కనుగొనడంలో మంచివి.
ఒక security researcher అంతకంటే కష్టమైనది చేస్తారు. వారు exposed entry points ను గుర్తిస్తారు, untrusted data ను application గుండా follow చేస్తారు, trust boundaries ను ప్రశ్నిస్తారు, మరియు అనేక code pieces కలిసి ఒక నిజమైన vulnerability గా ఎలా మారవచ్చనే theory ను నిర్మిస్తారు.
GitHub Security Lab, ఆ investigation process లో కొంత భాగాన్ని reusable AI-assisted workflow గా మార్చగలమా అని experiment చేస్తోంది.
September 28, 2026 న, team తన open-source Taskflow Agent ను ఉపయోగించి Android applications లో 24 vulnerabilities ను కనుగొని report చేసినట్లు వివరించింది. ఈ సంఖ్య గమనించదగినది, కానీ మరింత ముఖ్యమైన పరిణామం investigation ఎలా structure చేయబడింది అనేది.
Researchers కేవలం ఒక LLM కి repository ఇచ్చి bugs కనుగొనమని అడగలేదు.
వారు model కి follow చేయడానికి ఒక security investigation process ఇచ్చారు.
Source code నుంచి vulnerability hypothesis వరకు
GitHub యొక్క system taskflows ను ఉపయోగిస్తుంది.
ఒక taskflow అనేది reusable sequence of instructions మరియు tools, ఇది AI agent కి ఒక నిర్దిష్ట security task ను step by step ఎలా చేయాలో చెబుతుంది.
Mobile applications కోసం, workflow మొదట application యొక్క exposed entry points గురించి information సేకరిస్తుంది. GitHub యొక్క ప్రస్తుత taskflows Android components, deep links, URL schemes, application extensions మరియు WebView JavaScript bridges ను గుర్తించగలవు, permissions మరియు export status వంటి information తో పాటు.
ఆ తరువాత workflow ఆ entry points ను పరిశోధించేటప్పుడు mobile-specific security guidance ను apply చేయగలదు.
Figure 1 కోసం accessible text alternative
ఏడు దశల sequence: 1. Source code, ఇక్కడే audit మొదలవుతుంది. 2. Attack surface ను discover చేయడం. 3. Exported activities, intents మరియు WebView bridges వంటి security entry points ను గుర్తించడం. 4. Android-specific knowledge ను apply చేయడం. 5. Vulnerability hypothesis ను develop చేయడం, ఇది ఒక candidate explanation, ఇంకా proof కాదు. 6. Behaviour నిజమా కాదా test చేయడం. 7. Researcher finding ను validate చేయడం. ముగింపు గమనిక model steps 2 నుంచి 5 వరకు అందిస్తుందని, steps 6 మరియు 7 కు ఇంకా technical మరియు human evidence అవసరమని పేర్కొంటుంది.
LLM ముఖ్యమైనదే, కానీ ఇది ఈ system లో ఒక భాగం మాత్రమే.
మరో ముఖ్యమైన భాగం దాని చుట్టూ ఉన్న investigation method.
Security expertise లో కొంత భాగం workflow లోకి మారింది
Traditional security tools లో కూడా expert knowledge ఉంటుంది.
ఒక static-analysis rule attacker-controlled input ఒక dangerous function కు చేరడాన్ని చూడవచ్చు. ఒక CodeQL query data ఒక source నుంచి sensitive sink కు కదలడాన్ని model చేయగలదు. ఒక secret scanner కు నిర్దిష్ట credentials ఏ forms లో ఉండవచ్చో తెలుసు.
AI-assisted auditing కొంత security knowledge ను వేరే విధంగా express చేయడానికి అనుమతిస్తుంది.
ప్రతి idea ను ఒక deterministic query గా encode చేయడానికి బదులు, ఒక researcher workflow కి externally reachable components ను గుర్తించమని, attacker ఏ input ను control చేయగలడో అర్థం చేసుకోమని, relevant vulnerability classes ను పరిశోధించమని, code అంతటా relationships ను follow చేయమని, మరియు apparently separate behaviours ను కలిపి చూడగలరో లేదో నిర్ధారించమని చెప్పగలరు.
AI model ఆ steps మధ్య flexible reasoning ను అందిస్తుంది.
ఇది ఒక ముఖ్యమైన తేడాకు దారితీస్తుంది:
ఈ capability కేవలం better model నుంచి రాదు. ఇది model కి better investigation process ఇవ్వడం నుంచి కూడా వస్తుంది.
అందుకే ఈ taskflows కేవలం LLM కి source code review చేయమని అడగడం కంటే ఎక్కువ ఆసక్తికరమైనవి.
Security expertise మాయమవలేదు. దాంట్లో కొంత భాగం model follow చేసే workflow లోకి packaged అయింది.
ఒక నిజమైన ఉదాహరణ: OsmAnd లో trust boundary ను follow చేయడం
GitHub ఇచ్చిన ఉదాహరణల్లో ఒకటి Android navigation application OsmAnd కు సంబంధించినది.
Researchers ఒక exported Android activity ను కనుగొన్నారు, ఇది trusted internal path నుంచి రావాల్సిన intent data ను accept చేసేది.
ఆ component externally accessible గా ఉండడం వల్ల, మరో application ఆ values ను అందించగలిగేది.
తరువాత investigation ఆ data కు ఏమి జరిగిందో follow చేసింది.
GitHub reported ప్రకారం, ఒక attacker imported settings ను manipulate చేయగలడు, network configuration ను మార్చగలడు, మరియు map-tile లేదా routing requests ను attacker-controlled infrastructure కు redirect చేయగలడు. ఇది location information ను expose చేయవచ్చు.
Figure 2 కోసం accessible text alternative
ఆరు దశల chain: 1. External application, ఫోన్లో ఇప్పటికే ఉన్న మరే app అయినా. 2. Exported Android activity, permission లేకుండా చేరగల component. 3. Attacker-controlled intent data, internal గా ఉండాల్సిన values. 4. Application settings మార్చబడ్డాయి, expected check లేకుండా import చేయబడ్డాయి. 5. Network requests redirect అయ్యాయి, map-tile మరియు routing traffic తరలించబడింది. 6. Location information attacker-controlled infrastructure కు expose అయింది. ముగింపు గమనిక risk ఏదైనా ఒక suspicious code line నుంచి కాకుండా, chain నుంచి వస్తుందని పేర్కొంటుంది.
ఇది కేవలం ఒక dangerous function ను గుర్తించడం కాదు.
ఈ security problem, ఒక component ను ఎవరు call చేయగలరు, వారు ఏ information ను control చేయగలరు, ఆ information application state ను ఎలా మారుస్తుంది, మరియు ఆ మార్పు వల్ల తరువాత ఏమి జరుగుతుంది అనేది అర్థం చేసుకోవడం నుంచి పుడుతుంది.
ఆ రకమైన relationship, AI-assisted code investigation attention పొందడానికి ఒక కారణం.
ఇది ఒక AI system ఈ రకమైన ప్రతి vulnerability ను కనుగొంటుందని నిరూపించదు. అలాగే existing security techniques అదే issue ను కనుగొనలేవని కూడా నిరూపించదు.
ఇది flexible code reasoning ఒక AppSec workflow కి మరో ఉపయోగకరమైన layer ను ఎక్కడ జోడించగలదో చూపిస్తుంది.
Candidate finding అనేది evidence కాదు
ఇక్కడే limitations కూడా ముఖ్యమవుతాయి.
GitHub reports ప్రకారం, తమ models ఒక suspected vulnerability నిజంగా exploitable అవునా కాదా అనేది తప్పుగా అర్థం చేసుకోవచ్చు. అవి mitigating behaviour ను overlook చేయవచ్చు, low-impact issues ను return చేయవచ్చు, మరియు severity ను తప్పుగా estimate చేయవచ్చు.
ఒక code path dangerous గా కనిపిస్తూనే, runtime లో harmless గా ఉండవచ్చు.
ఉదాహరణకు, attacker-controlled data ఎక్కడో write చేయబడగలదని ఒక model చూసి, application లోని మరో భాగం దాన్ని consume చేస్తుందని నిర్ధారించవచ్చు. కానీ తరువాత జరిగే runtime behaviour దానికి బదులు trusted data కు ప్రాధాన్యత ఇవ్వవచ్చు.
Model యొక్క explanation ఒప్పించేలా ఉండవచ్చు.
అయినా ఆ vulnerability ఇంకా ఉండకపోవచ్చు.
Security teams కోసం, ఇది ఒక అవసరమైన boundary ను సృష్టిస్తుంది:
AI ఒక vulnerability hypothesis ను generate చేయగలదు. Exploitability కు ఇంకా evidence అవసరం.
ఆ evidence ఒక proof of concept, runtime testing, instrumentation, ఒక debugger, fuzzing, లేదా manual reproduction నుంచి రావచ్చు.
Claimed vulnerability ఎంత serious అయితే, ఈ తేడా అంత ముఖ్యమవుతుంది.
GitHub యొక్క ముందటి results ఆ gap పరిమాణాన్ని చూపిస్తాయి
GitHub Security Lab March 2026 లో మరో ఉపయోగకరమైన experiment ను publish చేసింది.
Researchers సంబంధిత auditing taskflows ను 40 కంటే ఎక్కువ repositories పై run చేశారు, ప్రధానంగా multi-user web applications పై.
LLM మొదట్లో 1,003 possible issues ను suggest చేసింది.
మరో audit stage, deduplication మరియు filtering తరువాత, researchers 91 candidates ను manually inspect చేశారు.
వారు తిరస్కరించారు:
- 20 ను false positives గా.
- 52 ను report చేయడానికి తగినంత low-impact గా.
వారు report చేయదగినంత serious గా భావించిన 19 findings ను retain చేశారు.
Figure 3 కోసం accessible text alternative
1,003 potential AI-generated issues నుంచి 91 manually investigated candidates వరకు funnel. ఒక branch 20 false positives మరియు 52 low-impact findings తొలగించబడ్డాయని చూపిస్తుంది. Funnel 19 reported vulnerabilities వద్ద ముగుస్తుంది. ముగింపు గమనిక raw AI output విజయానికి తప్పు కొలమానం అని, మరియు ఎన్ని candidates technical verification ను తట్టుకుని fix చేయడానికి తగినవిగా ఉన్నాయో అనేదే ముఖ్యమని పేర్కొంటుంది.
ఈ సంఖ్యలు ఉపయోగకరమైనవి ఎందుకంటే raw AI output ఎందుకు విజయానికి తప్పు కొలమానమో అవి చూపిస్తాయి.
ఒక agent వందల findings ఇచ్చినంత మాత్రాన AppSec team కు ప్రయోజనం కలగదు.
ఉపయోగకరమైన ప్రశ్న ఇది: serious technical verification ను ఎన్ని తట్టుకుంటాయి, మరియు fix చేయడానికి ఎన్ని తగినంత ముఖ్యమైనవి?
ఇది existing AppSec tools పక్కన ఎక్కడ సరిపోతుంది?
AI-assisted auditing ను ప్రతి existing security technique కి replacement గా treat చేయకూడదు.
వేర్వేరు tools వేర్వేరు రకాల evidence ను అందిస్తాయి.
| Capability | ఇది దేనిలో మంచిది |
|---|---|
| Static analysis | తెలిసిన code patterns మరియు data flows కోసం repeatable checks |
| Fuzzing | Generated inputs తో నిజమైన software ను exercise చేసి runtime behaviour ను గమనించడం |
| AI-assisted investigation | Code relationships ను explore చేసి vulnerability hypotheses ను develop చేయడం |
| Security researcher | Exploitability, severity, context మరియు business impact ను validate చేయడం |
ఈ capabilities కలిసి పనిచేయగలవు.
GitHub యొక్క September 2026 fuzzing work ఈ విధానాన్ని వివరిస్తుంది. దాని taskflow ఒక LLM agent ను ఉపయోగించి promising targets ను గుర్తిస్తుంది, build systems ను అర్థం చేసుకుంటుంది, fuzzing harnesses ను సృష్టిస్తుంది, coverage ను review చేస్తుంది మరియు crashes ను triage చేస్తుంది, అయితే AFL++ fuzzing ను నిజంగా చేసే execution engine గానే ఉంటుంది.
ఒక LLM AppSec toolchain మొత్తాన్ని replace చేస్తుందని ఆశించడం కంటే ఇది మరింత realistic దిశ.
AI ఎక్కడ మరియు ఎలా investigate చేయాలో నిర్ణయించడంలో సహాయపడగలదు.
Deterministic tools తాము test చేయడానికి design చేయబడినదాన్ని test చేయగలవు.
ఆ evidence మీద ఆధారపడిన security judgement బాధ్యత మనుషులదే.
Taskflow ఒక security engineering asset గా మారుతుంది
సంస్థలు ఇలాంటి workflows ను క్రమం తప్పకుండా ఉపయోగించడం మొదలుపెడితే మరో పరిణామం ఉంటుంది.
Taskflow తనకు తానే engineering discipline అవసరమయ్యే విషయంగా మారుతుంది.
Model, prompt, available tools లేదా investigation sequence ను మార్చడం results ను మార్చవచ్చు.
అందుకే ఒక mature implementation కనీసం ఐదు విషయాల గురించి ఆలోచించాల్సి ఉంటుంది.
Versioning
ఏ model మరియు workflow version ఒక finding ను ఇచ్చిందో teams కు తెలియాలి.
Regression testing
Known test cases, మార్చిన workflow గతంలో కనుగొన్న problems ను ఇప్పటికీ detect చేస్తుందా అని సమాధానం ఇవ్వడంలో సహాయపడతాయి.
Model-change testing
కొత్త model కి మారడాన్ని automatic గా improvement గా భావించకూడదు. Security behaviour ను మళ్లీ test చేయాలి.
Permissions
Source code ను మాత్రమే చదివే agent, software ను compile చేయడానికి, commands execute చేయడానికి లేదా proofs of concept run చేయడానికి అనుమతించబడిన agent కంటే వేరే risk ను కలిగి ఉంటుంది.
Evidence
ఒక security team, ఒక ముఖ్యమైన finding ఎలా ఉత్పత్తి అయిందో మరియు దాన్ని ఏ evidence confirm చేసిందో trace చేయగలగాలి.
Agent tools execute చేయగలిగినప్పుడు ఇది ముఖ్యంగా ముఖ్యమవుతుంది.
GitHub తన auditing taskflows ను ఒక sandboxed environment లో run చేయాలని సిఫారసు చేస్తుంది, మరియు Taskflow Agent documentation దాని Docker image ను తనకు తానే ఒక security boundary గా భావించకూడదని హెచ్చరిస్తుంది.
ఒక AI security agent code run చేయగలిగిన తరువాత, security design లో agent తనను తాను కూడా చేర్చుకోవాలి.
AI ఎక్కడ సహాయపడగలదు, ఎక్కడ ఇంకా evidence అవసరం
| Stage | AI సహాయపడగలదా | ఇంకా దేనికి verification అవసరం |
|---|---|---|
| Repository reconnaissance | అవును | ముఖ్యమైన omissions |
| Code-path investigation | అవును | నిజమైన behaviour గురించి assumptions |
| Vulnerability hypothesis | అవును | Reproduction |
| Proof-of-concept creation | అవును | Safe execution మరియు result validation |
| Exploitability | Support మాత్రమే | Runtime evidence |
| Severity | Support మాత్రమే | Security judgement |
| Business impact | పరిమిత context | Human decision |
AI “security లో మంచిదేనా” అని అడగడం కంటే ఈ విభజన ఎక్కువ ఉపయోగకరమైనది.
వేర్వేరు stages కు వేర్వేరు స్థాయిల confidence అవసరం.
AppSec teams ఇప్పుడు ఏమి చేయాలి?
ఈ experiments ఆసక్తికరమైన results ఇస్తున్నందున existing security tools ను replace చేయాల్సిన అవసరం లేదా autonomous vulnerability-hunting programme launch చేయాల్సిన అవసరం లేదు.
Narrow గా మొదలుపెట్టడం ఒక మంచి విధానం.
ఒక investigation task తో మొదలుపెట్టండి
Measurable అయిన దానికి AI assistance ను ఉపయోగించండి: entry-point discovery, existing scanner findings యొక్క analysis, authorization-path investigation, fuzzing-harness creation, లేదా మరో well-defined task.
పూర్తి ఫలితాన్ని కొలవండి
Raw AI findings ను విజయంగా లెక్కించకండి.
ఎన్ని findings reproduce చేయగలరో, ఎన్ని అర్థవంతమైనవిగా నిరూపించబడతాయో, researcher work ఎంత మిగులుతుందో, మరియు workflow ఏ ముఖ్యమైన problems ను miss చేస్తుందో కొలవండి.
Evidence gate ను కొనసాగించండి
ఒక AI-generated finding, దాని explanation ఒప్పించేలా అనిపించినంత మాత్రాన confirmed vulnerability గా మారకూడదు.
Exploitability, severity మరియు impact కు ఇంకా తగిన technical evidence అవసరం.
Headline కంటే workflow ఎక్కువ ముఖ్యం
GitHub యొక్క 24 Android vulnerabilities, AI-assisted security investigation మరింత capable గా మారుతోందనడానికి ఉపయోగకరమైన evidence.
కానీ మరింత ముఖ్యమైన పరిణామం వాటి వెనుక ఉన్న workflow.
Security researchers తమ investigation methods లో కొంత భాగాన్ని AI models చుట్టూ reusable processes గా package చేయడం మొదలుపెడుతున్నారు.
Model flexible reasoning ను అందిస్తుంది. Taskflow structure మరియు security knowledge ను అందిస్తుంది. Existing security tools deterministic మరియు runtime evidence ను అందిస్తాయి. ఫలితంగా వచ్చిన vulnerability story నిజమా కాదా నిర్ణయించే బాధ్యత human researchers దే.
అది AppSec లో AI కోసం మరింత practical దిశను సూచిస్తుంది:
మంచి security investigation methods ను repeatable workflows గా మార్చండి, flexible reasoning సహాయపడే చోట AI ను ఉపయోగించండి, మరియు result ను నమ్మడానికి ముందు evidence ను కోరండి.
ఇది ఒక AI model కి bugs కనుగొనమని అడగడం కంటే ఎక్కువ demanding.
ఇది security engineering నిజంగా ఎలా పనిచేస్తుందో దానికి చాలా దగ్గరగా ఉంటుంది.
References and further reading
- How we found 24 Android vulnerabilities using our open source AI security agent: GitHub Security Lab, September 28, 2026 న publish చేయబడింది. Android audit workflow, vulnerability examples, మరియు exploitability, severity assessment చుట్టూ ఉన్న limitations కోసం primary evidence.
- How to scan for vulnerabilities with GitHub Security Lab’s open source AI-powered framework: GitHub Security Lab, March 2026 న publish చేయబడింది. 40 కంటే ఎక్కువ repositories అంతటా విస్తృత experiment ను మరియు 1,003 candidate issues నుంచి 19 reported vulnerabilities వరకు progression ను documents చేస్తుంది.
- SecLab Taskflows: GitHub Security Lab. Mobile entry-point discovery తో సహా, reusable security workflows కోసం open-source implementation మరియు documentation.
- AI-powered fuzzing with the GitHub Security Lab Taskflow Agent: GitHub Security Lab, September 24, 2026 న publish చేయబడింది. AFL++ నిజమైన runtime fuzzing చేస్తుండగా, LLM-driven taskflows fuzzing workflow లోని కొన్ని భాగాలను ఎలా coordinate చేయగలవో చూపిస్తుంది.
