GitHub’s open source AI agent finds 24 Android flaws, including a Wikipedia account takeover
The GitHub Security Lab Taskflow Agent found 24 real Android vulnerabilities, from location tracking in OsmAnd to a full Wikipedia account takeover. AI can now find high-impact logic bugs, but still struggles to rank their severity.
September 28, 2026. Kevin Stubbings, of GitHub Security Lab, publishes the results of an audit run with the Taskflow Agent, their open source AI security agent: 24 Android vulnerabilities found and reported. September 28, 2026. The post details two critical cases, including location tracking in OsmAnd, a navigation app with over 10 million downloads. September 28, 2026. The same technique yields a Wikipedia account takeover through the official app. Why it matters: AI has crossed a threshold — it no longer finds only generic bug classes, but logic bugs with critical impact.
An agent that breaks the audit into guided steps
The Taskflow Agent rests on a simple principle: large models read code well, but get lost on big repositories. GitHub Security Lab’s answer is to split the search into taskflows — YAML files carrying prompts that spell out, step by step, what the model should look for.
For Android, two additions make the difference. A gather_mobile_entry_point_info.yaml taskflow isolates entry points — the places where attacker-controlled data can enter — and separates mobile from non-mobile ones, so the model targets the right attack surface. A second file, classify_application_local.yaml, injects a list of mobile vulnerability classes (confused deputy, insecure broadcasts, and so on) that the model must check at each entry point.
The benefit is twofold. The strict prompt and repeated passes keep obvious bugs from slipping through, while the broad prompt lets the model exercise its creativity on the odd cases. The whole thing runs with one command:
git clone https://github.com/GitHubSecurityLab/seclab-taskflows
cd seclab-taskflows
./scripts/audit/run_mobile.sh myorg/myrepo The cost is explicit and documented: a GitHub Copilot license, premium model requests, and token consumption that can climb fast. A full audit takes one to two hours on a medium-sized repository.
OsmAnd: tracking a user through their map tiles
The first example shows the impact goes well beyond a crash. OsmAnd exports an activity called MapActivity, which handles settings-file imports through intent extras: settings_version, silent_import, replace, export_type_list_key.
The problem is structural. That activity is exported, so any app can send it an intent with arbitrary extras — Android provides no mechanism to restrict what an external caller can set. By injecting silent_import and replace, an app with no permissions at all imports a configuration without the user noticing.
The imported configuration contains the URL template for map tiles. OsmAnd formats each tile’s URL like this:
MessageFormat.format(urlTemplate, zoom, x, y) By replacing that template with an attacker-controlled domain, every loaded tile reports its x, y coordinates and zoom level back to the server. The attacker then reconstructs the exact coordinates of every tile displayed, plus the origin and destination of every computed route — all invisible to the user. Full location tracking, obtained through a simple configuration manipulation.
Wikipedia: account takeover via deeplink
The second example is graver still. The Wikipedia app registers a hook for the wikipedia:// deeplink, meant to open only pages on the Wikipedia domain. A logic bug in the hostname parser allows loading any URL.
The cascade is devastating. The attacker makes the app open a page ending in wikipedia.org — for example evil-wikipedia.org — which passes the domain check even though it is hostile. The page runs inside the app’s WebView, with arbitrary JavaScript, in a context normally considered safe. A second weakness, in SharedPreferenceCookieManager.kt, checks the cookie domain with a simple endsWith:
if (domain.endsWith(domainSpec)) {
buildCookieList(cookieList, cookiesForDomainSpec, null)
} Chaining the two flaws, the attacker harvests the victim’s long-lived cookies: their username, session token and long-lived token, valid across all Wikimedia projects — every Wikipedia, Commons, Wikidata, Meta. A complete account takeover, from a single click on a link.
What AI can do — and what it cannot
The verdict is nuanced, which is what makes it credible. GitHub Security Lab notes that the models excel at finding vulnerabilities — to the point of surfacing low-severity bugs of little consequence — and hold fine-grained knowledge of security-relevant APIs, able to produce proofs of concept that need almost no modification.
The weak spot lies elsewhere: severity estimation. The model routinely over- or under-rates real impact, blind to mitigating factors — a path traversal confined to external storage, or internal data that overwrites attacker-controlled data. Hence the team’s golden rule: every finding must be reviewed by a human researcher who knows mobile.
The distribution of findings is telling. Many were simple bugs such as path traversal, but a handful were critical, and the classes cluster exactly where a mobile researcher would look: cross-app scripting in a WebView and exposed JavaScript bridges. Android’s platform hardening is strong, so the residual risk concentrates in the seams between apps — intents, deeplinks and exported components.
The 24 findings were reported through the GitHub Security Lab advisories page, which lists disclosures as they land — a public trail that doubles as a benchmark for what AI-assisted audits can surface.
The lesson fits in one sentence: AI has become a research amplifier, not a replacement. It multiplies the bugs found, including high-impact logic bugs; it does not decide, alone, what deserves a fix.
What it changes for defense
The same tooling shifts the game on both sides. For a defender, these taskflows turn a weeks-long manual audit into a few hours — the time to set up the repository, run the script and triage the results. For an attacker, the reverse logic applies: a model able to produce near-exploitable proofs of concept lowers the cost of discovery.
The asymmetry is not new. It is about velocity: when discovery becomes automated, the window between publishing an app and exploiting its flaws shrinks. The answer is not to flee AI, but to fold it into the development cycle — auditing before release, not after compromise.
Verdict
If you maintain an open source Android app, running the Taskflow Agent taskflows against your repository is a high-yield move: the cost is a Copilot license and a few hours, the gain is a triage of logic bugs that classical scanners miss. If you consume third-party code or apps, take away the threshold crossed — attackers now have the same tooling, and the line between audit and exploitation becomes a question of who pulls the trigger first. In both cases, keep a human in the loop for severity: AI finds, the human decides.