AI-Assisted AppSec: Security Expertise को Repeatable Workflows में बदलना

GitHub Security Lab ने Android vulnerabilities की पहचान के लिए संरचित AI workflows का उपयोग किया। इससे अधिक महत्वपूर्ण सबक यह है कि security expertise, AI reasoning, deterministic tools और human verification एक साथ कैसे काम कर सकते हैं।

इन भाषाओं में पढ़ें: English · తెలుగు · हिन्दी

Diagram showing source code passing through an AI-assisted security investigation and evidence validation before becoming a confirmed vulnerability.

GitHub Security Lab के Android रिसर्च से पता चलता है कि संरचित taskflows AI को कमजोरियों की जांच में किस तरह मदद कर सकते हैं, और साथ ही यह भी कि candidate findings को अब भी तकनीकी साक्ष्य और मानवीय निर्णय की आवश्यकता क्यों है।

पारंपरिक security scanner उन patterns को पहचानने में अच्छे होते हैं जिन्हें पहचानने के लिए उन्हें डिज़ाइन किया गया है।

एक security researcher इससे कहीं कठिन काम करता है। वे उजागर entry points की पहचान करते हैं, untrusted data को application में follow करते हैं, trust boundaries पर सवाल उठाते हैं, और यह theory बनाते हैं कि कोड के कई हिस्से मिलकर एक वास्तविक vulnerability कैसे बन सकते हैं।

GitHub Security Lab यह प्रयोग कर रही है कि क्या इस investigation process का कुछ हिस्सा एक reusable AI-assisted workflow में बदला जा सकता है।

28 सितंबर 2026 को टीम ने बताया कि उसने अपने open-source Taskflow Agent का उपयोग करके Android applications में 24 vulnerabilities खोजीं और रिपोर्ट कीं। यह संख्या उल्लेखनीय है, लेकिन अधिक महत्वपूर्ण विकास यह है कि investigation को किस तरह संरचित किया गया था।

Researchers ने केवल किसी LLM को एक repository देकर bugs खोजने को नहीं कहा।

उन्होंने model को एक security investigation process दिया, जिसका उसे पालन करना था।

सोर्स कोड से एक vulnerability hypothesis तक

GitHub का सिस्टम taskflows का उपयोग करता है।

एक taskflow निर्देशों और tools का एक reusable क्रम है जो AI agent को बताता है कि किसी विशेष security task को चरण दर चरण कैसे पूरा करना है।

Mobile applications के लिए, workflow सबसे पहले application के उजागर entry points के बारे में जानकारी इकट्ठा करता है। GitHub के मौजूदा taskflows Android components, deep links, URL schemes, application extensions और WebView JavaScript bridges की पहचान कर सकते हैं, साथ ही permissions और export status जैसी जानकारी भी।

इसके बाद workflow उन entry points की जांच करते समय mobile-specific security guidance लागू कर सकता है।

From source code to a validated finding Seven-step sequence from source code through attack-surface discovery, entry-point identification, Android-specific investigation and a vulnerability hypothesis to technical testing and researcher validation. {“creator”:”TechiesJournal”,”author”:”Prasad Kukkala”,”asset”:”ai-assisted-appsec-figure-1”,”source_revision”:”ai-assisted-appsec-v1-2026-09-30”,”created”:”2026-09-30”,”rights”:”Copyright 2026 TechiesJournal. All rights reserved.”} 1 Source code Where the audit begins 2 Discover the attack surface Map what is externally reachable 3 Identify security entry points Exported activities, intents, WebView bridges 4 Apply Android-specific knowledge Mobile security patterns and mitigations 5 Develop a vulnerability hypothesis A candidate explanation, not yet proof 6 Test whether the behaviour is real Technical validation of the hypothesis 7 Researcher validates the finding Human judgement confirms the result The model provides steps 2 to 5. Steps 6 and 7 still require technical and human evidence. TECHIESJOURNAL
चित्र 1: एक संरचित AI-assisted audit attack-surface discovery, investigation, hypothesis generation और technical validation को अलग-अलग रखता है, बजाय इसके कि model के output को एक confirmed finding मान लिया जाए।
चित्र 1 के लिए टेक्स्ट विवरण

सात चरणों वाला क्रम: 1. Source code, जहां audit शुरू होता है। 2. Attack surface की खोज करें। 3. Security entry points की पहचान करें, जैसे exported activities, intents और WebView bridges। 4. Android-specific knowledge लागू करें। 5. एक vulnerability hypothesis विकसित करें, जो एक candidate explanation है, अभी proof नहीं। 6. जांचें कि behaviour वास्तविक है या नहीं। 7. Researcher finding को validate करता है। एक closing note बताता है कि model steps 2 से 5 प्रदान करता है, जबकि steps 6 और 7 के लिए अब भी technical और human evidence चाहिए।

LLM महत्वपूर्ण है, लेकिन यह इस सिस्टम का केवल एक हिस्सा है।

दूसरा महत्वपूर्ण हिस्सा वह investigation method है जो इसके इर्द-गिर्द बनाया गया है।

Security expertise का कुछ हिस्सा अब workflow में समा गया है

पारंपरिक security tools में भी expert knowledge होता है।

एक static-analysis rule attacker-controlled input को किसी dangerous function तक पहुंचते हुए ढूंढ सकता है। एक CodeQL query data के source से किसी sensitive sink तक जाने को model कर सकता है। एक secret scanner जानता है कि विशेष credentials किन रूपों में दिख सकते हैं।

AI-assisted auditing कुछ security knowledge को अलग ढंग से व्यक्त करने की अनुमति देता है।

हर idea को एक deterministic query के रूप में encode करने के बजाय, एक researcher workflow से कह सकता है कि वह externally reachable components की पहचान करे, यह समझे कि attacker किस input को control कर सकता है, प्रासंगिक vulnerability classes की जांच करे, कोड में संबंधों का पीछा करे, और यह तय करे कि क्या अलग-अलग दिखने वाले behaviours को मिलाया जा सकता है।

AI model इन चरणों के बीच लचीला reasoning प्रदान करता है।

इससे एक महत्वपूर्ण अंतर सामने आता है:

यह capability केवल एक बेहतर model से नहीं आती। यह model को एक बेहतर investigation process देने से भी आती है।

यही कारण है कि ये taskflows किसी LLM से केवल सोर्स कोड की समीक्षा करवाने से कहीं अधिक दिलचस्प हैं।

Security expertise गायब नहीं हुई है। इसका कुछ हिस्सा उस workflow में पैक कर दिया गया है, जिसका model पालन करता है।

एक वास्तविक उदाहरण: OsmAnd में एक trust boundary का पीछा करना

GitHub के उदाहरणों में से एक Android navigation application OsmAnd से जुड़ा था।

Researchers को एक exported Android activity मिली जो ऐसा intent data स्वीकार करती थी, जिसके बारे में यह अपेक्षा थी कि वह एक trusted internal path से आएगा।

क्योंकि यह component externally accessible था, इसलिए कोई अन्य application ये values प्रदान कर सकता था।

इसके बाद investigation ने यह जांचा कि उस data का आगे क्या हुआ।

GitHub ने बताया कि एक attacker imported settings में बदलाव कर सकता था, network configuration बदल सकता था और map-tile या routing requests को attacker-controlled infrastructure की ओर redirect कर सकता था। इससे location information उजागर हो सकती थी।

Following the OsmAnd trust boundary Six-step sequence from an external application through an exported Android activity, attacker-controlled intent data, changed application settings and redirected network requests to exposed location information. {“creator”:”TechiesJournal”,”author”:”Prasad Kukkala”,”asset”:”ai-assisted-appsec-figure-2”,”source_revision”:”ai-assisted-appsec-v1-2026-09-30”,”created”:”2026-09-30”,”rights”:”Copyright 2026 TechiesJournal. All rights reserved.”} 1 External application Any other app already on the phone 2 Exported Android activity A component reachable without permission 3 Attacker-controlled intent data Values meant to stay internal 4 Application settings changed Imported without the expected check 5 Network requests redirected Map-tile and routing traffic moved 6 Location information exposed To attacker-controlled infrastructure The risk comes from the chain, not from any single suspicious line of code. TECHIESJOURNAL
चित्र 2: रिपोर्ट किया गया OsmAnd issue attacker-controlled data को application के कई behaviours में पीछा करने पर निर्भर था, न कि केवल एक संदिग्ध API की पहचान करने पर।
चित्र 2 के लिए टेक्स्ट विवरण

छह चरणों वाली श्रृंखला: 1. External application, यानी फोन पर पहले से मौजूद कोई भी अन्य app। 2. Exported Android activity, यानी बिना permission के पहुंच योग्य एक component। 3. Attacker-controlled intent data, यानी ऐसे values जिन्हें internal रहना था। 4. Application settings बदले गए, बिना अपेक्षित जांच के import किए गए। 5. Network requests redirect किए गए, map-tile और routing traffic को स्थानांतरित किया गया। 6. Location information attacker-controlled infrastructure के सामने उजागर हुई। एक closing note बताता है कि यह जोखिम इस chain से आता है, न कि किसी एक संदिग्ध लाइन ऑफ कोड से।

यह केवल एक dangerous function को पहचानने का मामला नहीं है।

Security problem इस समझ से उभरती है कि कौन किसी component को call कर सकता है, वे किस जानकारी को control कर सकते हैं, वह जानकारी application state को कैसे बदलती है, और उस बदलाव के कारण बाद में क्या होता है।

यही तरह का संबंध एक कारण है कि AI-assisted code investigation ध्यान आकर्षित कर रही है।

यह यह साबित नहीं करता कि AI system इस तरह की हर vulnerability खोज लेगा। यह भी साबित नहीं करता कि मौजूदा security techniques वही issue नहीं खोज सकती थीं।

यह यह दिखाता है कि लचीला code reasoning कहां एक AppSec workflow में एक और उपयोगी परत जोड़ सकता है।

एक candidate finding evidence नहीं है

यहीं पर limitations भी महत्वपूर्ण हो जाती हैं।

GitHub बताता है कि उसके models यह गलत समझ सकते हैं कि कोई संदिग्ध vulnerability वास्तव में exploitable है या नहीं। वे mitigating behaviour को नज़रअंदाज़ कर सकते हैं, low-impact issues लौटा सकते हैं, और severity का गलत अनुमान लगा सकते हैं।

एक code path देखने में खतरनाक लग सकता है, जबकि runtime पर वह हानिरहित हो।

उदाहरण के लिए, एक model यह देख सकता है कि attacker-controlled data कहीं लिखा जा सकता है और यह निष्कर्ष निकाल सकता है कि application का कोई अन्य हिस्सा उसे उपयोग करेगा। लेकिन बाद का runtime behaviour trusted data को प्राथमिकता दे सकता है।

Model की व्याख्या convincing लग सकती है।

फिर भी vulnerability मौजूद न हो, ऐसा हो सकता है।

Security teams के लिए यह एक आवश्यक सीमा बनाता है:

AI एक vulnerability hypothesis बना सकता है। Exploitability को अब भी evidence चाहिए।

वह evidence एक proof of concept, runtime testing, instrumentation, एक debugger, fuzzing, या manual reproduction से आ सकता है।

दावा की गई vulnerability जितनी गंभीर होगी, यह अंतर उतना ही अधिक महत्वपूर्ण हो जाता है।

GitHub के पहले के परिणाम इस gap का आकार दिखाते हैं

GitHub Security Lab ने मार्च 2026 में एक और उपयोगी प्रयोग प्रकाशित किया था।

Researchers ने संबंधित auditing taskflows को 40 से अधिक repositories पर चलाया, जो मुख्यतः multi-user web applications थीं।

LLM ने शुरुआत में 1,003 संभावित issues सुझाए।

एक और audit चरण, deduplication और filtering के बाद, researchers ने 91 candidates की मैन्युअल जांच की।

उन्होंने अस्वीकार किए:

  • 20 को false positives के रूप में।
  • 52 को रिपोर्ट करने लायक बहुत कम impact वाला मानते हुए।

उन्होंने 19 findings रखे, जिन्हें उन्होंने रिपोर्ट करने योग्य पर्याप्त गंभीर माना।

From AI suggestion to reportable vulnerability Funnel from 1,003 AI-generated candidate issues to 91 manually investigated, with 20 false positives and 52 low-impact findings removed, ending at 19 reported vulnerabilities. {“creator”:”TechiesJournal”,”author”:”Prasad Kukkala”,”asset”:”ai-assisted-appsec-figure-3”,”source_revision”:”ai-assisted-appsec-v1-2026-09-30”,”created”:”2026-09-30”,”rights”:”Copyright 2026 TechiesJournal. All rights reserved.”} 1,003 Potential issues suggested by the AI audit 91 Manually investigated by researchers 20 false positives 52 too low-impact 19 Reported vulnerabilities Raw AI output is the wrong measure of success. What matters is how many candidates survive technical verification and are worth fixing. TECHIESJOURNAL
चित्र 3: GitHub का पहले वाला audit प्रयोग शुरुआती AI-generated candidates और रिपोर्ट करने लायक मानी गई vulnerabilities के बीच के बड़े gap को दिखाता है।
चित्र 3 के लिए टेक्स्ट विवरण

1,003 संभावित AI-generated issues से लेकर मैन्युअल रूप से जांचे गए 91 candidates तक का funnel। एक शाखा दिखाती है कि 20 false positives और 52 low-impact findings हटा दिए गए। यह funnel 19 रिपोर्ट की गई vulnerabilities पर समाप्त होता है। एक closing note बताता है कि raw AI output सफलता का गलत पैमाना है, और असली सवाल यह है कि कितने candidates technical verification में टिकते हैं और उन्हें ठीक करना उचित है।

ये संख्याएं इसलिए उपयोगी हैं क्योंकि वे दिखाती हैं कि raw AI output सफलता का गलत पैमाना क्यों है।

एक AppSec team को सिर्फ इसलिए फायदा नहीं होता क्योंकि कोई agent सैकड़ों findings पैदा करता है।

असली सवाल यह है: कितने गंभीर technical verification में टिकते हैं और उन्हें ठीक करना इतना महत्वपूर्ण है?

यह मौजूदा AppSec tools के साथ कहां फिट बैठता है?

AI-assisted auditing को हर मौजूदा security technique का replacement नहीं माना जाना चाहिए।

अलग-अलग tools अलग-अलग तरह का evidence प्रदान करते हैं।

अलग-अलग AppSec क्षमताएं कहां evidence प्रदान करती हैं
क्षमतायह किस काम में सबसे अच्छी है
Static analysisज्ञात code patterns और data flows की दोहराई जाने योग्य जांच
Fuzzingवास्तविक software को generated inputs से चलाना और runtime behaviour देखना
AI-assisted investigationCode संबंधों की खोज और vulnerability hypotheses विकसित करना
Security researcherExploitability, severity, context और business impact को validate करना

ये क्षमताएं एक साथ काम कर सकती हैं।

GitHub का सितंबर 2026 का fuzzing कार्य इसी approach को दिखाता है। इसका taskflow एक LLM agent का उपयोग यह पहचानने के लिए करता है कि कौन-से targets आशाजनक हैं, build systems को समझने, fuzzing harnesses बनाने, coverage की समीक्षा करने और crashes को triage करने के लिए, जबकि AFL++ ही वह execution engine बना रहता है जो असली fuzzing करता है।

यह किसी LLM से पूरे AppSec toolchain को बदलवाने की अपेक्षा करने से कहीं अधिक व्यावहारिक दिशा है।

AI यह तय करने में मदद कर सकता है कि कहां और कैसे जांच की जाए।

Deterministic tools वही टेस्ट कर सकते हैं जिसके लिए उन्हें डिज़ाइन किया गया है।

इस evidence के आधार पर बनने वाले security judgement की जिम्मेदारी अब भी मनुष्यों की ही रहती है।

Taskflow खुद एक security engineering asset बन जाता है

अगर organisations इस तरह के workflows का नियमित उपयोग शुरू करते हैं, तो इसका एक और परिणाम सामने आता है।

Taskflow खुद एक ऐसी चीज़ बन जाता है जिसे engineering discipline की आवश्यकता होती है।

Model, prompt, उपलब्ध tools या investigation sequence में बदलाव परिणामों को बदल सकता है।

इसलिए एक परिपक्व implementation को कम से कम पांच बातों के बारे में सोचने की आवश्यकता होगी।

Versioning

Teams को यह पता होना चाहिए कि किस model और workflow version ने कोई finding उत्पन्न की।

Regression testing

ज्ञात test cases यह जवाब देने में मदद कर सकते हैं कि बदला हुआ workflow अब भी उन समस्याओं को पकड़ता है, जिन्हें उसने पहले पकड़ा था।

Model-change testing

किसी नए model पर जाने को स्वतः ही एक सुधार नहीं माना जाना चाहिए। Security behaviour को फिर से test करने की आवश्यकता है।

Permissions

एक agent जो केवल सोर्स कोड पढ़ता है, उस agent से अलग जोखिम प्रस्तुत करता है जिसे software compile करने, commands execute करने या proofs of concept चलाने की अनुमति है।

Evidence

एक security team को यह पता लगा पाना चाहिए कि कोई महत्वपूर्ण finding कैसे उत्पन्न हुई और किस evidence ने उसकी पुष्टि की।

यह विशेष रूप से महत्वपूर्ण हो जाता है जब agent tools को execute कर सकता है।

GitHub अनुशंसा करता है कि उसके auditing taskflows को एक sandboxed environment में चलाया जाए, और Taskflow Agent documentation चेतावनी देता है कि उसकी Docker image को खुद एक security boundary नहीं माना जाना चाहिए।

एक बार जब कोई AI security agent code चला सकता है, तो security design में उस agent को भी शामिल करना ज़रूरी हो जाता है।

AI कहां मदद कर सकता है, और कहां अब भी evidence की आवश्यकता है

AppSec में AI कहां मदद करता है, और अब भी किस चीज़ के verification की आवश्यकता है
चरणक्या AI मदद कर सकता हैअब भी किसे verify करने की आवश्यकता है
Repository reconnaissanceहांमहत्वपूर्ण चूकें
Code-path investigationहांवास्तविक behaviour के बारे में assumptions
Vulnerability hypothesisहांReproduction
Proof-of-concept creationहांSafe execution और result validation
Exploitabilityकेवल सहायकRuntime evidence
Severityकेवल सहायकSecurity judgement
Business impactसीमित contextHuman decision

यह विभाजन इस सवाल से कहीं अधिक उपयोगी है कि AI सिर्फ “security में अच्छा” है या नहीं।

अलग-अलग चरणों के लिए विश्वास के अलग-अलग स्तर चाहिए।

AppSec teams को अभी क्या करना चाहिए?

केवल इसलिए मौजूदा security tools को बदलने या एक autonomous vulnerability-hunting programme शुरू करने की आवश्यकता नहीं है क्योंकि ये प्रयोग दिलचस्प परिणाम दे रहे हैं।

एक बेहतर approach यह है कि छोटे दायरे से शुरुआत की जाए।

एक investigation task से शुरुआत करें

AI सहायता का उपयोग किसी measurable चीज़ के लिए करें: entry-point discovery, मौजूदा scanner findings का विश्लेषण, authorization-path investigation, fuzzing-harness creation, या किसी अन्य well-defined task के लिए।

पूरे परिणाम को मापें

Raw AI findings को सफलता के रूप में न गिनें।

यह मापें कि कितनी findings reproduce की जा सकती हैं, कितनी सार्थक साबित होती हैं, researcher का कितना काम बच जाता है, और workflow किन महत्वपूर्ण समस्याओं को छोड़ देता है।

एक evidence gate बनाए रखें

एक AI-generated finding को केवल इसलिए एक confirmed vulnerability नहीं बन जाना चाहिए क्योंकि उसकी व्याख्या convincing लगती है।

Exploitability, severity और impact को अब भी उचित technical evidence की आवश्यकता है।

Workflow, headline से कहीं अधिक मायने रखता है

GitHub की 24 Android vulnerabilities इस बात का उपयोगी evidence हैं कि AI-assisted security investigation अधिक सक्षम होती जा रही है।

लेकिन इससे कहीं अधिक महत्वपूर्ण विकास इनके पीछे का workflow है।

Security researchers अपने investigation methods के हिस्सों को AI models के इर्द-गिर्द reusable processes में पैक करना शुरू कर रहे हैं।

Model लचीला reasoning प्रदान करता है। Taskflow structure और security knowledge प्रदान करता है। मौजूदा security tools deterministic और runtime evidence प्रदान करते हैं। यह तय करने की जिम्मेदारी अब भी human researchers की रहती है कि परिणामी vulnerability story वास्तव में सच है या नहीं।

इससे AppSec में AI के लिए एक अधिक व्यावहारिक दिशा का संकेत मिलता है:

अच्छे security investigation methods को repeatable workflows में बदलें, जहां लचीला reasoning मदद करे वहां AI का उपयोग करें, और परिणाम पर भरोसा करने से पहले evidence की मांग करें।

यह किसी AI model से bugs खोजने को कहने से कहीं अधिक मांग वाला काम है।

यह इस बात के भी कहीं अधिक करीब है कि security engineering वास्तव में कैसे काम करती है।

References and further reading

  • How we found 24 Android vulnerabilities using our open source AI security agent: GitHub Security Lab, 28 सितंबर 2026 को प्रकाशित। Android audit workflow, vulnerability examples और exploitability एवं severity assessment से जुड़ी सीमाओं का प्राथमिक evidence।
  • How to scan for vulnerabilities with GitHub Security Lab’s open source AI-powered framework: GitHub Security Lab, मार्च 2026 में प्रकाशित। 40 से अधिक repositories पर हुए व्यापक प्रयोग और 1,003 candidate issues से 19 रिपोर्ट की गई vulnerabilities तक की प्रगति का दस्तावेज़ीकरण करता है।
  • SecLab Taskflows: GitHub Security Lab। Reusable security workflows के लिए open-source implementation और documentation, जिसमें mobile entry-point discovery भी शामिल है।
  • AI-powered fuzzing with the GitHub Security Lab Taskflow Agent: GitHub Security Lab, 24 सितंबर 2026 को प्रकाशित। यह दिखाता है कि LLM-driven taskflows किस तरह fuzzing workflow के कुछ हिस्सों का समन्वय कर सकते हैं, जबकि असली runtime fuzzing AFL++ करता है।
सुधार बताएं

सुधार सीधे एडिटर तक पहुँचते हैं; ये अपने आप कभी प्रकाशित नहीं होते। किसी अकाउंट की ज़रूरत नहीं।