AppXpose AppXpose
Methodology

Last updated August 2026. Reflects v7.4.1.

How AppXpose
actually works.

Privacy claims are easy to make and hard to verify. This page documents the full pipeline: every step, every scoring factor, every data source, and every thing we don't touch. If something here is wrong, email mahere@appxpose.app and we will fix it.

I. The scan pipeline
01

You pick an app

AppXpose lists every app installed on your device via the Android PackageManager API. You tap one. Nothing happens until you do.

02

We unpack the APK locally

The PackageManager hands us the installed APK file path. We read its DEX bytecode and AndroidManifest in-process on Dispatchers.IO. Bytes never leave the phone.

03

Pattern-match against known tracker signatures

Class names and package paths are compared against a curated signature set (sourced from Exodus Privacy, our own additions, and community-discovered patterns). Transitive dependencies from parent SDKs are deduplicated so the count reflects actual distinct trackers. This is deterministic pattern matching, not AI. See II. Detection methodology below.

04

Score deterministically, then AI-analyze

The local results are sent (HMAC-signed, tied to a device fingerprint, no name or account) to our Cloudflare edge. A deterministic pre-score is calculated from multiple risk factors. Then an LLM generates the final score breakdown, the developer profile, the paywall analysis, and the natural-language explanations. See III. Risk scoring model.

05

You see a score and a story

Not just a number. What was found, why it matters, who made the app, where the data goes, and what changed since the last scan. You can save the report, share it, or vote on it for the community.

06

Everything lands in the App Index

Every scan result is stored in a local database — risk scores, tracker lists, permission audits, all timestamped. The App Index lets you browse, search, and compare your installed apps in one place. Filter by risk level, sort by last scan date, spot which apps got worse over time. It turns individual scans into a structured record of what is running on your device.

II. Detection methodology

How we find trackers in bytecode.

Tracker detection in AppXpose is deterministic pattern matching, not AI. The LLM never decides whether a tracker is present. A signature either matches or it doesn't.

Signature sources

SourceExodus Privacy
Count136
MethodPackage path prefixes from their public database (ODbL licensed)
SourceCommunity discovery
Count137
MethodManual APK analysis + auto-confirmed from user scans via community discovery pipeline. When 3+ distinct packages contain the same unknown SDK, it becomes a confirmed signature.
SourceTotal
Count270+
MethodGrows automatically from community scans — no app update required

Match types

  • →Package path prefix: e.g. com.facebook.appevents, com.adjust.sdk. If a class in the DEX starts with this path, it's a match. Each signature is compiled to a regex at app start.
  • →Prefix deduplication: when a parent SDK (e.g. com.google.android.gms.ads) matches, sub-packages like com.google.android.gms.ads.doubleclick are not counted separately. Known transitive dependencies (e.g. Firebase Analytics bundled by AdMob) are suppressed when the parent is confirmed present.
  • →Suspicious class detection: classes that look like SDKs but don't match any known signature are flagged and sent to the server for community-driven discovery (see below).

Community tracker discovery

Every scan contributes to growing the signature database. When the DEX scanner finds classes that look like third-party SDKs but don't match any known signature, they are reported anonymously to the server. Once the same unknown class prefix appears in 3 or more distinct packages from different devices, it is auto-confirmed as a new tracker signature. Confirmed signatures sync daily to all devices via a background worker — no app update required. This means the detection engine gets better with every scan the community runs.

Note: estimated tracker information provided in the AI analysis section uses LLM knowledge and is always marked as "(estimated)". Only DEX-verified trackers are reported without a caveat.

III. Risk scoring model

Two-phase scoring: deterministic, then AI.

Phase 1: Deterministic pre-score

Before the LLM sees any data, a deterministic scoring function runs on the server. It evaluates multiple risk signals including permission analysis (context-aware across 45 Play Store categories), update history, installation source, APK integrity, breach history, tracker presence, signing certificate verification, and known-malware database matches. Positive signals (clean history, no trackers, recent updates) reduce the score.

The pre-score produces a baseline from 0 to 100. Risk levels: LOW (0-29), MEDIUM (30-59), HIGH (60-79), CRITICAL (80-100). The exact weights and thresholds are part of our proprietary scoring engine and will be open-sourced with the detection engine.

Context-aware scoring

Permissions that are normal for one category are suspicious in another. A weather app asking for location scores differently than a flashlight asking for the same thing. Our scoring engine maps expected permissions across 45 Play Store categories and adjusts risk accordingly. This mapping is continuously refined by a data-driven baseline that learns from real scan data.

Phase 2: AI-assisted analysis

The pre-score and all raw data (permissions, trackers, breach status, app metadata) are sent to an LLM which generates the user-facing breakdown. The LLM produces six category scores visible in the app:

  1. Known Data Collection Practices
  2. Parent Company & Geopolitical Risk
  3. Estimated Trackers & Ad Networks
  4. Encryption & Data Transit
  5. Update Frequency & Patch Responsiveness
  6. Regulatory Compliance & Transparency

The LLM also generates: the developer profile (company, server locations, GDPR posture), the paywall and monetization analysis, the data-sharing probability estimates, and the natural-language explanations next to every finding. All AI-estimated data is explicitly marked as "(estimated)" in the app.

IV. What the AI does and doesn't do

AI generates

  • → Risk score breakdown (6 categories)
  • → Paywall and monetization analysis
  • → Developer profile (company, servers, GDPR)
  • → Data-sharing probability estimates
  • → Natural-language explanations
  • → All marked "(estimated)" where applicable

AI does not

  • → Detect trackers (that's pattern matching)
  • → Check breaches (that's HIBP)
  • → List permissions (that's the Android API)
  • → Read your apps' content or data
  • → Access your device outside the scan
  • → Make the final tracker count (DEX does)
V. Data sources
SourceAndroid PackageManager
Used forAPK file access, metadata, permissions
Leaves device?No
SourceDEX bytecode reader
Used forTracker signature matching
Leaves device?No
SourceExodus Privacy DB
Used for136 tracker signatures
Leaves device?Cached on our edge. Phone sends package name only.
SourceHave I Been Pwned
Used forDeveloper breach history
Leaves device?Nothing about you. Our server fetches the public breach list and matches it against app vendors.
SourceLLM (Anthropic)
Used forScore breakdown, descriptions, developer profiling
Leaves device?Receives app metadata + permissions. No bytecode, no personal data.
SourceCommunity votes
Used forCrowd-sourced risk signal
Leaves device?Stored anonymously on Cloudflare D1. No author identity.
SourceMalwareBazaar (abuse.ch)
Used forKnown-malware hash lookup
Leaves device?APK SHA256 checked server-side. No personal data sent. Open community DB.
SourceAppXpose CertNet (TOFU + F-Droid)
Used forKnown-good signing cert verification
Leaves device?Crowd-sourced signing cert DB. Self-populates via trust-on-first-use from real scans. Seeded with F-Droid verified certs.
SourceGoogle Play Integrity
Used forVerifying the app was installed from Google Play (observation mode, no blocking)
Leaves device?Integrity token sent to Google's servers for verification. Contains app + device attestation. No personal data.
SourceGoogle AdMob (free tier only)
Used forRewarded video ads (opt-in, never banners or interstitials). SDK is not initialized for premium / GUARD users.
Leaves device?When a free user voluntarily watches a rewarded ad: device model, OS version, app version, language/region, IP address, Google Advertising ID (unless reset in Android settings), ad-interaction events. Standard AdMob SDK behavior.
VI. What we collect, and what we don't

We never collect

  • Email addresses or accounts
  • Google Advertising ID (AppXpose itself never reads it; AdMob may read it for free users who voluntarily watch a rewarded ad)
  • IMEI, phone number, SIM data
  • Location, contacts, photos
  • App content or messages
  • Persistent identifiers across reinstalls

We do collect

  • Device fingerprint hash (rotates on reinstall)
  • Quota counter (5 scans/week free tier)
  • Anonymous community votes (no author identity)
  • Scan raw data for ML training (package name, permissions, tracker matches, suspicious classes — no personal data, used to improve detection accuracy)
  • APK signing certificate hashes (for CertNet trust-on-first-use verification)
VII. Architecture
--- Your Phone ---
AppXpose (Kotlin + Jetpack Compose) — EN · DE · ES · PT
> DEX bytecode reader (runs on Dispatchers.IO)
> Tracker signature matcher (270+ patterns, grows via community sync)
> Local cache (Room DB, v10 keys)
> SharedPrefs: scan history flag
> GUARD Workers (5 alerts, daily via WorkManager)
> Play Integrity (async warmup, observation mode)
| HMAC-SHA256 signed, no auth, HTTPS
v
--- Cloudflare Edge ---
Worker + D1 database
> Deterministic pre-score (multiple factors)
> LLM API call (score breakdown + descriptions)
> Cached signatures (Exodus, 5-min warm, EN)
> Community tracker discovery (auto-confirm at 3+ packages)
> Quota counter (per device fingerprint)
> HIBP public breach list (matched server-side)
> Community votes (anonymous)

Everything between phone and edge is HTTPS plus HMAC-SHA256. The HMAC key is a server-side secret that can be rotated independently of app updates, with a grace period where old and new keys are accepted simultaneously. The "device fingerprint" is a one-way hash of stable hardware characteristics. It cannot be reversed into a person, and it resets on reinstall or factory reset.

VIII. Source code status

Closed source. Here's why, and what changes when.

AppXpose is a solo project that is still in heavy development. The scan pipeline was rewritten multiple times in the last months as new detection approaches surfaced. Publishing the code today would mean publishing a moving target that breaks by next week. That is not useful to anyone.

The plan: once the detection engine's interfaces stabilize (the signature format, the scoring model inputs, the pipeline boundaries), the engine will be opened. Not the whole app, but the part that matters: the component that reads bytecode and decides what a tracker is. That is the part the community should be able to audit, challenge, and improve.

Detection engine (signature matcher, DEX reader)
WILL BE OPENED
Scoring model (pre-score factors, weights)
Published with the detection engine (proprietary until open-source release)
App UI (Kotlin + Compose)
Stays closed (UI code is not security-relevant)
Backend (Cloudflare Worker)
Stays closed (infrastructure orchestration, business logic)

In the meantime, this page documents everything the code does and how. The code is closed. The methodology is not. If you find a gap in this documentation, email us and we will close it.

Found something we missed?

Security claims should be falsifiable. If you see a gap in this methodology, a factor we are not accounting for, or a claim that does not match reality, tell us. We will fix it and credit you in the next release notes.

mahere@appxpose.app