Review 4 of 6: Scoring model, calibration, false positives, scam domain knowledge

Reviewer lens: does the two-ledger log-odds + decision-table design produce verdicts that are right, explainable, and stable? Are the reason codes grounded in how US smishing actually looks in 2024–2026?

Companion artifacts (all in this folder):


0. Bottom line

The architecture is right: two ledgers, reason codes with published weights, and a decision table instead of a single threshold. Keep it. It has five defects that will cause real misclassifications, and each one is fixable in a line or two of spec:

  1. Correlated codes are double-counted. A toll smish fires brand, urgency, fee, payment, link and TLD codes all at once, and naive summing treats correlated evidence as independent. Fix: group codes into families with caps, or take the max within a family.
  2. The prior must depend on sender kind. As drafted, UNKNOWN_SENDER is both the base rate and a weighted risk code, so it counts twice. Fix: pick the prior by sender kind (contact / short code / toll-free / long code / email gateway / international), and make UNKNOWN_SENDER weight 0, kept only for the explanation.
  3. Claims in the message text can reach Likely legitimate. Under draft rule 5, a scammer typing "This is Carlos from your building" earns context points. Fix: Likely legitimate requires at least one hard (attacker-uncontrollable) code.
  4. Gate order hides tripwires. "completeness < 40 → Cannot assess" runs first, so a visible "send me the code" in a half-redacted preview gets swallowed. Fix: tripwires run on whatever text is visible, before the completeness gate.
  5. Several tripwires are false-positive-prone as worded. Three examples:

Fix: add negation guards, require imperative-plus-object, and restrict the brand table to high-impersonation brands.

Everything below makes these concrete.


1. Critique of the draft scoring design

1.1 The two ledgers: keep, but separate attacker-controlled evidence from attacker-uncontrollable evidence

The ledgers split "who is this" from "what is this asking". That split is correct and it is the product's core differentiator. The missing axis is who controls the evidence:

Attacker-controlled (cheap to fake)Not attacker-controlled (hard evidence)
Self-identification text ("this is X from Y")Sender is in Contacts (Person URI is a contacts lookup)
Brand names in textUser marked the sender trusted
"New number" claimsUser previously sent a substantive reply in this thread
Relationship claims ("we met at…")Short code in the bundled brand table
Using the user's first name (it is in data-broker dumps)Self-ID matches a specific entry in the private context ledger
Calendar match
Every link resolves to the named brand's allowlisted eTLD+1
Sender has ≥3 prior clean messages over ≥7 days

Rule: soft codes may move a message within Unrecognized. Only hard codes can lift it to Likely legitimate. In the reference implementation this is the hard flag. Corpus pair G22 vs G23 tests it: identical Carlos text, with and without a ledger entry.

Two subtleties:

1.2 Log-odds summing: correlated codes must be capped by family

Naive-Bayes summing assumes the codes are conditionally independent. They are not. "late fee", "$6.99", "pay by 10/04" and "toll" all come from the same scam template. Fix:

FamilyMembersAggregation
LINKURL_*, LOOKALIKE_DOMAIN, BRAND_DOMAIN_MISMATCHsum, cap +3.0
ASKOTP / gift card / crypto / remote-access / credential / payment / callbacksum, cap +3.0
PRETEXTtoll, package, bank, job, govt, carrier, crypto-invest, prizemax (one pretext per message)
PRESSUREurgency, penalty, small feesum, cap +1.0
SENDERbrand-from-long-code, intl-claims-US-entity, group blastsum, cap +1.0
SOCIALopener, wrong-number pivot, platform shift, new number, family emergency, secrecysum, cap +2.0
EVASIONreply-to-enable-link, injection, homoglyphsum, cap +2.5

The caps keep a single template from saturating the score through repetition. Independent families still add. A toll smish with a lookalike link and the reply-Y trick should be at about 100, and is (G16, L = +8.5).

The prior carries the base rate by sender kind:

sender_kindprior logitscam_risk with no codesRationale
contact / user-trusted−4.02Compromise is possible but rare. A separate gate handles it (§1.5).
short code (5–6 digits)−2.58Carrier-vetted (CTIA short code program). Scam short codes exist but are uncommon.
toll-free (8XX)−2.012Verification is required for toll-free A2P, but enforcement is weaker.
US long code−1.518Mixes personal numbers and 10DLC business numbers.
international−0.831Smishing Triad uses foreign SIMs. Real friends abroad also text.
email-to-SMS gateway−0.538Heavily abused by toll/USPS kits. A few legitimate uses (legacy alerts, people texting from email).
sender unavailable−1.518Also lowers completeness.

All values are starting points (§5). G38 (a +44 friend) and G39 (a bare "Hello, are you free?" from an email gateway) pin the requirement that no sender kind alone reaches Suspicious (40). G39 sits at 38 on purpose: if calibration pushes the email-gateway prior above about −0.4, that test fails loudly, and that is the intended alarm.

Important: the sidecar sees a post-filter distribution. Google Messages' own spam filter suppresses notifications for what it catches, so the sidecar sees what Google missed plus all ham. The priors therefore should be lower than "fraction of unknown texts that are scams", and public datasets (full of easy spam) overstate how separable the residual is.

1.3 Are the five dimensions coherent? Is legitimacy_confidence redundant?

As drafted, legitimacy = identity + context − risk is redundant. It is a mix of three other displayed numbers, and once risk is folded in it becomes a noisy second copy of 100 − scam_risk. That is the exact conflation the PRD forbids ("unknown is not malicious").

The PRD requires five dimensions, so do not drop it. Redefine it so it carries information the others do not:

The point of keeping legitimacy and risk as separate quantities: both can be high at once, and that is the most informative state the product can show. Two examples:

A single score cannot say either sentence.

UI consequence for contacts. For contacts, scam_risk with a −4 prior stays low (G24 reports 23) while the verdict says Suspicious. That will confuse users. Either display a content risk computed with the neutral long-code prior (−1.5) next to sender trust, or show only the triggered codes for contacts. Recommended: always compute both scam_risk (sender-aware, used for gating) and content_risk (prior −1.5, for display and for the compromise gate).

1.4 How analysis_completeness should interact

The draft checks completeness first, which hides visible tripwires. Use this ordering instead:

  1. If no text is available (hidden preview, OS OTP redaction, "New message"), completeness is below 40 and the verdict is Cannot assess. No tripwire can fire without text, so nothing is hidden.
  2. Partial text: run every detector on the visible text. Tripwires and risk codes count fully, because presence is evidence. Absence of risk codes is weaker evidence when completeness is below 80:

G33 ("Your account has been tempor…") must land on Unrecognized, not Likely legitimate.

  1. Interruption policy keys on identity_confidence, not on the verdict. G34 is Mom with a hidden preview: verdict Cannot assess, identity 95. In "Quiet unknowns" mode Mom must still alert. If interruption keys on the verdict, the product silences the user's mother whenever lock-screen privacy is on.

1.5 Revised decision table (as implemented in tools/score.py)

no visible text (completeness < 40)            -> Cannot assess
sender is contact / user-trusted:
    tripwire OR content risk-code sum >= 2.5   -> Suspicious + POSSIBLE_ACCOUNT_COMPROMISE
                                                  (Likely scam/phishing if scam_risk >= 70)
    else                                       -> Recognized
any tripwire                                   -> Likely scam/phishing if scam_risk >= 70, else Suspicious
scam_risk >= 70                                -> Likely scam/phishing
scam_risk >= 40                                -> Suspicious
(identity >= 60 OR context >= 60) AND >=1 hard code AND risk < 40
                                               -> Likely legitimate
otherwise                                      -> Unrecognized / insufficient context

Logit thresholds: 40 ↔ L = −0.405 and 70 ↔ L = +0.847.

Differences from the draft:

1.6 Other specific problems


2. Reason-code catalog (starter, ~60 codes)

Text normalization runs first and is shared by every detector:

  1. Apply NFKC.
  2. Record and strip zero-width characters (U+200B–U+200D, U+2060, U+FEFF).
  3. Map confusables (Cyrillic/Greek → Latin) via the Unicode confusables table. Record whether anything changed.
  4. Lowercase a working copy.
  5. De-fang: hxxp→http, [.]/(.)/ dot →., and usps . com → usps.com.

URL extraction:

Columns: weight is in log-odds. T = tripwire. H = hard evidence (attacker cannot control).

2.1 Sender-kind (prior selector; explanation-only codes)

CodeRulePriorFP notes
UNKNOWN_SENDERNo contacts-lookup URI in android.people.list and not in the local trust listweight 0 (prior carries it)Never decisive, by construction.
(kind) short codeSender is ^\d{5,6}$−2.5Some 3–4 digit carrier codes (e.g. 728, 611): treat ^\d{3,6}$ as short code.
(kind) toll-free`^\+?1?(800\833\844\855\866\877\888)\d{7}$`−2.0
(kind) long code^\+?1?[2-9]\d{2}[2-9]\d{6}$−1.510DLC business numbers live here too.
(kind) internationalE.164 with country code ≠ +1, or NANP non-US (+1 876/809/… Caribbean)−0.8Caribbean +1 area codes are a classic one-ring/"wangiri" vector. They look domestic, so keep a list.
(kind) email gatewaySender contains @, or is a known gateway pseudo-number−0.5Some real alerts (old school/county systems) send via email-to-SMS. Hence a prior, not a tripwire.

2.2 Identity ledger

CodeWtHRuleFP / abuse notes
CONTACT_MATCH+5.0HPerson.uri starts with content://com.android.contacts/contacts/lookup/, or the android.title display name is non-numeric with a lookup URI. Verify in the spike.Spoofed caller-ID SMS is rare but possible. Hence the compromise gate.
USER_TRUSTED+5.0HHMAC(normalized number) is in the local trust listBehaves like a contact.
THREAD_USER_REPLIED+2.5Handroid.messages tail contains a message with sender == null/self whose text has ≥3 words or ≥15 characters and is not in the trivial-reply setPig-butchering victims also reply substantively. That is why this is identity, not a risk reducer.
THREAD_USER_REPLIED_TRIVIAL0The user's only reply matches `^\s*(y\yes\n\no\1\ok\k\stop\wrong number\who is this\??)\s*$`Exists to block reply-Y laundering.
SHORT_CODE+1.0HSender kind is short code
KNOWN_SHORT_CODE_BRAND+2.0 id, −1.0 riskHShort code is in the bundled {code → brand} table and the brand named in the text (if any) equals that brandThe table must come from brand-published sources. Short codes are country-specific.
SENDER_HISTORY_CLEAN+1.0HPer-sender state: ≥3 prior messages, first seen ≥7 days ago, no risk code ≥1.0 everSlow-burn pig-butchering builds clean history. Cap at +1.0 for that reason.
SELF_IDENTIFIED+0.7`\b(this is\it'?s\i'?m\my name is)\s+([A-Z][a-z]+(\s[A-Z][a-z]+)?)(\s+(from\with\at\of)\s+([A-Z][\w&'.-]+(\s[A-Z][\w&'.-]+){0,4}))?, or a leading ^[A-Z][\w &'.-]{2,40}:\s` brand prefixScammers self-identify constantly (G19, G26). Soft by design.
LINK_DOMAIN_VERIFIED+1.5 id, −1.0 riskHAll URLs' eTLD+1 are in the allowlist of the brand named in the text, or match a domain the user attached to a ledger entryA real-brand link can be paired with a callback-number vishing ask. CALLBACK_NUMBER still fires.

2.3 Context ledger

CodeWtHRuleFP notes
EXPECTED_SENDER_SPECIFIC+2.5HThe self-ID entity, or a ≥2-token proper-noun span, fuzzy-matches (Jaro-Winkler ≥0.92) a named non-expired ledger entry ("Bright Smile Dental", "Carlos" + "Ashford")Common first names alone ("Carlos") must not match. Require entity + org, or a full name.
EXPECTED_SENDER_GENERIC+1.0Text names a carrier/brand category that matches a ledger entry ("FedEx", "UPS", "Amazon")Soft: scammers send carrier lures by the million. Never lowers risk (G09).
CALENDAR_MATCH+1.5HOrg name or attendee in text matches an event within ±3 days (optional READ_CALENDAR)Calendar invites from scammers exist. Require the org to match the event title, not just the description.
THREAD_CONTINUITY+1.5MessagingStyle tail has ≥1 prior exchange in both directions, and the new message is topically continuous (shared nouns/time references)Same reply-Y caveat. Requires a non-trivial user reply.
APPOINTMENT_REMINDER+0.8`\b(appt\appointment\visit\reservation\booking)\b and a date/time \b(mon\tue\wed\thu\fri\sat\sun\\d{1,2}/\d{1,2})\b.*\b\d{1,2}(:\d{2})?\s?(am\pm)\b and reply\s+(c\y\1\confirm)`Scam "appointment" lures exist (fake DMV appointments). Soft weight.
OTP_DELIVERY+0.5\b\d{4,8}\b near `(code\otp\passcode\pin) and a negation (do not\don'?t\never)\s+(share\give) or (will never ask)`Drop the digit span before storage. Usually redacted on Android 15+ anyway.
IN_PERSON_DELIVERY+0.5`(outside\at (the\your) (door\gate\lobby)\downstairs\which (building\unit\door))` and package/delivery wording, and no URL
ADDRESSES_USER_BY_NAME+0.3The user's configured first name appears in the textData-broker leaks make this cheap. Keep it tiny.

2.4 Risk ledger: LINK family (sum, cap +3.0)

CodeWtTRuleFP notes
URL_PRESENT+0.3≥1 URL extractedReal reminders contain links. Tiny on purpose. Merchant descriptors are not links: a scheme-less domain right after at/from in a purchase/charge/transaction context ("purchase at BESTBUY.COM") is skipped. Without that rule, every real bank alert for an online purchase trips BRAND_DOMAIN_MISMATCH (found by running the corpus, G06).
URL_SHORTENER+0.7eTLD+1 in {bit.ly, tinyurl.com, t.co, is.gd, rb.gy, cutt.ly, ow.ly, shorturl.at, tiny.cc, s.id, t.ly, …}Brand-owned shorteners (e.g. amzn.to, vz.to): put them in the brand allowlist, not here.
URL_RISKY_TLD+1.0TLD in {top, xyz, icu, cyou, vip, shop, live, info, click, buzz, sbs, cfd, bond, online, lat, win, rest, mom, lol, help, support, today}.shop/.live/.info have real users. That is why it is not a tripwire. Refresh this list from SmishTank/Unit 42 TLD stats.
URL_IP_LITERAL+1.5Host matches ^\d{1,3}(\.\d{1,3}){3}$ or contains [ (IPv6)Almost no legitimate SMS uses IP literals.
URL_OBFUSCATED+0.8De-fanging changed something, or the URL is split across spaces/line breaks to defeat linkification
LOOKALIKE_DOMAIN+2.5TAny of: (a) xn-- punycode label; (b) a brand/agency token from the brand table appears in a label of the host but the eTLD+1 is not on that brand's allowlist (ezpass.com-tollpay.top, usps.redelivery-care.cyou, wellsfargo.secure-verify-login.com); (c) eTLD+1 second-level label within Damerau-Levenshtein ≤2 of a brand label (fedx.com, uspss.com) after confusable mapping(b) needs a brand table with care: "chase" in chasebank-realtors.com? Restrict tokens to ≥4 characters and to high-impersonation brands. Expect some FPs on fan/affiliate sites; the tripwire only forces Suspicious, not Likely scam.
BRAND_DOMAIN_MISMATCH+2.5TText names a brand from the high-impersonation table (USPS, UPS, FedEx, DHL, Amazon, Apple, Google, Microsoft, PayPal, Venmo, Zelle, Cash App, Chase, BofA, Wells Fargo, Citi, Capital One, Amex, Navy Federal, USAA, IRS, SSA, DMV, E-ZPass, SunPass, FasTrak, TxTag, Peach Pass, Toll By Plate, Netflix, T-Mobile, Verizon, AT&T, Coinbase, …), and some URL's eTLD+1 is not on that brand's allowlistExcluded on purpose: small businesses, clinics and dentists, which routinely use vendor domains (Weave, Solutionreach, NexHealth, Klara, Podium; G04). Marketing click-trackers (e.g. click.e.brand.com) usually share the eTLD+1; if not, add them to the allowlist.

2.5 Risk ledger: ASK family (sum, cap +3.0)

CodeWtTRuleFP notes
OTP_SHARE_REQUEST+3.0T`(send\give\tell\forward\text\read)\s+(me\us\it\back)?.{0,20}\b(the\that\your\a)?\s*(\d[- ]?digit\s+)?(code\otp\pin\passcode\verification) or (what('?s\is) the code\accidentally sent .{0,30}code). Negation guard: suppress if within 40 characters of (do not\don'?t\never\will never ask\no one from)`Real 2FA texts say "never share" (G11, must_not_include).
GIFT_CARD_PAYMENT+3.0TImperative buy/pay verb `(buy\purchase\pick up\get me\pay (with\in\using)) within 40 characters of (gift\s?cards?\itunes\apple card\google play\steam\ebay card\vanilla\razer gold), or (send\scratch).{0,30}(codes?\pins?\numbers on the back)` together with gift-card wordsRetail promos ("get a $10 gift card when you buy 2") must not match. Require the card to be the payment instrument (G27).
CRYPTO_PAYMENT_DEMAND+3.0T`(send\deposit\transfer\pay\top up) within 40 characters of (btc\bitcoin\usdt\tether\eth\crypto\wallet address\bitcoin atm), or a BTC/ETH/TRON address regex \b(bc1\[13])[a-km-zA-HJ-NP-Z1-9]{25,39}\b\\b0x[a-fA-F0-9]{40}\b\\bT[1-9A-HJ-NP-Za-km-z]{33}\b`Users who trade crypto get real exchange notices. Those come from known short codes and don't demand deposits.
REMOTE_ACCESS_APP+3.0T`(install\download\open) within 40 characters of (anydesk\teamviewer\quicksupport\rustdesk\screenconnect\ultraviewer\zoho assist)`Legitimate IT support by SMS is vanishingly rare for consumers.
CREDENTIAL_REQUEST+1.5`(verify\confirm\update\validate\restore)\s+(your\s+)?(identity\account\card\details\information\login\ssn\social security)`Real banks say "log in to the app to review". Hence not a tripwire.
PAYMENT_REQUEST+1.0`\b(pay\payment\settle\balance due\outstanding (amount\balance)\fee)\b` with an imperative or a deadlineReal bills (G31) carry it. Neutralized by LINK_DOMAIN_VERIFIED + KNOWN_SHORT_CODE_BRAND.
CALLBACK_NUMBER+0.8A phone number in the text plus `(call\contact)` in an alert/fraud contextReal fraud alerts say "call the number on the back of your card", which has no number. Explicit numbers in alerts are a vishing tell.
REPLY_YES_NO_ALERT+0.3`reply\s+(yes\y)\s+or\s+(no\n)` in a transaction-alert contextReal banks do exactly this (G06). Tiny weight; the sender kind does the work.

2.6 Risk ledger: PRETEXT (max of family), PRESSURE (cap 1.0), SENDER (cap 1.0)

CodeWtTRuleFP notes / domain note
PRETEXT_TOLL+1.5`(toll\e-?z ?pass\sunpass\fastrak\txtag\peach pass\ipass\i-pass\ezdrive\good to go\toll by plate\turnpike) with (unpaid\outstanding\balance\due\overdue)`IC3 logged >60k toll complaints in 2024. Some agencies do text account holders who opt in. Those come from agency short codes with agency domains, and allowlisting handles them.
PRETEXT_PACKAGE_PROBLEM+1.0`(package\parcel\shipment\delivery) with (unable\failed\could not\incomplete\invalid\update (your )?address\redeliver\on hold\held\suspended\customs\fee)`Real "on hold" notices exist, but come from carrier short codes with carrier links. FTC's #1 text scam of 2024.
PRETEXT_BANK_ALERT+1.0`(suspicious\unusual\unauthori[sz]ed)\s+(activity\transaction\login\charge\sign-?in), did you (attempt\authori[sz]e\make), account (has been )?(locked\suspended\restricted\tempor)`FTC top-5 (fake fraud alerts). Smishing Triad pivoted to banks in 2025.
PRETEXT_JOB_TASK+1.5`\$\s?\d{2,4}\s?(-\s?\$?\d{2,4})?\s?(/\per\a\daily\each)\s?(day\daily\hour) or (part-?time\remote).{0,40}(flexible\no experience\\d+\s?min) or (app optimi[sz]ation\product boost\data optimi[sz]ation\rating tasks\click farm\merchant tasks)`Real recruiters rarely state day rates by text (G20 must_not). Task scams ≈ 40% of 2024 job-scam reports (FTC).
PRETEXT_GOVT+1.5`(irs\internal revenue\social security\ssa\dmv\department of motor vehicles\court\warrant\jury duty\stimulus\tax refund\administrative code)`Real jury/court SMS systems exist in some counties. They don't demand payment via link.
PRETEXT_CARRIER+1.0Carrier brand + `(reward\points\gift\thank you\expire\suspend\sim\port)`Real carrier bill notices (G31) lack reward/suspend words.
PRETEXT_CRYPTO_INVEST+1.5`(crypto\bitcoin\usdt\trading\invest(ment)?\forex\gold futures\signals\mentor\(\d+)x\b\guaranteed (return\profit))`Real friends talk about stocks. Single-message weight is moderate; the trajectory adds more.
PRETEXT_PRIZE_REFUND+1.0`(you('ve\have) won\winner\claim your\reward\refund (of\is ready)\free gift\selected)`Real loyalty texts from short codes. The sender prior absorbs them.
URGENCY+0.5`(immediately\asap\urgent\within \d+\s?(h\hrs?\hours?)\final (notice\reminder)\last chance\act now\pay now\expires? (today\soon)\by \d{1,2}/\d{1,2})`Standalone "today" is deliberately excluded: "delivery today" and "are you working today" are ham.
PENALTY_THREAT+0.7`(late fee\penalt\suspend\revok\legal action\arrest\collections\credit score\forfeit)`
SMALL_FEE+0.5`\$\s?(0?\.\d{2}\[1-9](\.\d{2})?\1[0-9](\.\d{2})?)\b` with fee/toll/redelivery$1.99 redelivery and $3–15 toll amounts are the templates.
BRAND_FROM_LONG_CODE+0.8A high-impersonation brand is named and the sender kind is long code or email gatewaySome banks and brands use 10DLC long codes legitimately. It also fires on a driver saying "your Amazon package" (G21, risk 33, still Unrecognized). Moderate weight; never escalates alone.
INTL_CLAIMS_US_ENTITY+1.0Sender kind international and the text names a US agency or brandReal US brands don't text US customers from +63/+44.
GROUP_BLAST+0.5isGroupConversation and ≥5 participants, none in contactsNeighborhood groups. Small weight.

2.7 Risk ledger: SOCIAL (cap 2.0) and EVASION (cap 2.5)

CodeWtTRuleFP notes
GENERIC_OPENER0Whole message ≤12 words and matches `^(hi\hey\hello)?[,!]?\s*(are you (working\free\available\around\there)\is this \w+\how are you\long time no see\you there)\b(\s+\w+){0,3}\s*[?!.]?` (0–3 trailing words: "…working today?", "…free this weekend?")Weight 0. It only arms the trajectory detector for 14 days. PRD golden case (G01).
MISADDRESSED_GREETING+0.3`^(hi\hey)?,?\s*is this (\w+)\??` and the name ≠ the user's configured nameReal misdials happen. Tiny.
NEW_NUMBER_CLAIM+0.3`(new (number\phone)\changed my number\lost my phone\dropped my phone\this is my new)`Real friends do this (G02). Becomes meaningful only combined with an ask.
NAME_MATCHES_OTHER_CONTACT+0.6Self-ID name equals a contact's display name, and that contact's number ≠ the sender (optional READ_CONTACTS; without it, the code is unavailable and explained as such)Must not raise identity (G03). Advice: "verify via Jake's saved number".
WRONG_NUMBER_PIVOT+1.0Trajectory: within 14 days of an opener or misaddressed greeting from the same sender, the message matches `(sorry\oops).{0,40}(wrong number\wrong person) with continued friendly conversation (nice to meet\fate\where are you from\my name is`)Real misdialers apologize and stop. The continuation is the tell.
PLATFORM_SHIFT+1.0`(whatsapp\telegram\signal\line app\wechat\kakao\viber\add me on\text me on)`Real colleagues say "WhatsApp me" too. Moderate weight.
INVESTMENT_PIVOT+1.5Trajectory: PRETEXT_CRYPTO_INVEST fires on a sender whose earlier messages were social/opener-only
FAMILY_EMERGENCY_MONEY+1.5Kinship term `(mom\mum\dad\grandma\grandpa\nana\auntie) with NEW_NUMBER_CLAIM or (it'?s me\can'?t talk)`, plus a money askUS analog of the UK/AU "Hi Mum" scam. Requires the money ask.
SECRECY+0.7`(don'?t tell\keep (this\it) between us\can'?t talk( right now)?\text only\in a meeting)`
REPLY_TO_ENABLE_LINK+2.5T`reply\s+(with\s+)?["']?(y\yes\1)["']? within 80 characters of (link\activate\re-?open\exit\copy.{0,20}(browser\safari\chrome))`Real "Reply Y to confirm your appointment" has no link-activation words, so it does not match. Documented 2025 iMessage-targeting trick, now in kit templates sent to Android too.
INJECTION_ATTEMPT+2.0T`(ignore (all\any\previous\prior)\s+(instructions\rules)\system (note\prompt\message)\s*:\(classify\mark\label) (this\me) as (safe\legit\trusted)\you are (an? )?(ai\assistant\model)\note to (ai\assistant\llm))`A friend writing "ignore previous instructions lol" would trip it. Acceptable: it only forces Suspicious unless other risk is present.
HOMOGLYPH_OR_ZERO_WIDTH+0.8Normalization changed brand-word characters (e.g. Cyrillic Ѕ in "UЅPЅ"), or zero-width characters were presentEmoji ZWJ sequences: exclude U+200D when it sits between emoji.

2.8 Meta / completeness codes (weight 0, always shown)

POSSIBLE_ACCOUNT_COMPROMISE (from the contact gate), PREVIEW_TRUNCATED, PREVIEW_HIDDEN (placeholder text such as "New message" or "Sensitive notification content hidden"), OTP_REDACTED_BY_OS, SENDER_ID_MISSING.


3. Golden corpus

The full corpus is in golden-corpus.json: 42 cases covering contact, US long code, short code, toll-free, email gateway and international senders.

Fictional numbers use 555-01XX and 07700 900XXX. Short code values for real brands are illustrative; the shipped table must come from brand-published lists.

Every case's expected codes and verdict are reproduced end-to-end by tools/detect.py + tools/score.py, starting from the raw text. If an implementer changes a weight or a regex, tools/corpus.py says which cases moved.

IDSenderGistLriskidctxExpected
G01long"Are you working today?"−1.5181212Unrecognized
G02long"Hey it's Jake, new number"−1.2232112Unrecognized
G03longsame; a contact "Jake" exists on another number−0.6352112Unrecognized (verify via saved number)
G04toll-freereal dentist reminder + Weave link, no ledger−1.7152123Unrecognized (low risk)
G05toll-freesame + ledger "Bright Smile Dental"−1.7152179Likely legitimate
G06shortChase "did you attempt $482.13? reply YES/NO"−2.2108512Likely legitimate
G07longsame + "call 1-888-…"+1.9872112Likely scam
G08longE-ZPass $6.99 / ezpass.com-tollpay.top+5.81001212Likely scam
G09longFedEx expected; link fedex-track.info+3.8982127Likely scam
G10shortreal FedEx with a fedex.com link−4.219627Likely legitimate
G11shortGoogle 2FA, "don't share"−3.537318Likely legitimate
G12shortOTP redacted by Android 15————Cannot assess
G13long"accidentally sent you a code, send it back"+2.0881212Likely scam
G14contactcontact asks for a code−1.0279512Suspicious + compromise
G15+63USPS redelivery + reply-Y+7.2100Likely scam
G16emailtoll + reply-Y+8.5100Likely scam
G17long"Is this David? This is Emma…"−1.2232112Unrecognized (MISADDRESSED_GREETING arms trajectory)
G18longpivot: fate / crypto / WhatsApp+2.088Likely scam
G19longtask scam, $300–800/day+1.073Likely scam
G20longreal recruiter by name−1.5182115Unrecognized
G21longdriver outside, no link ("Amazon" fires BRAND_FROM_LONG_CODE)−0.7331218Unrecognized
G22longCarlos @ The Ashford, ledger match−1.5182162Likely legitimate
G23longsame, no ledger−1.5182112Unrecognized
G24contactfriend shilling a crypto platform−1.2239512Suspicious + compromise
G25long"Mom it's me, new number, pay a bill"+1.582Likely scam
G26longCEO gift cards+3.597Likely scam
G27shortTarget promo "get a $10 gift card"−3.247712Likely legitimate
G28longIRS refund lookalike+4.899Likely scam
G29emailDMV "Administrative Code" ticket+6.8100Likely scam
G30longT-Mobile reward .shop+3.396Likely scam
G31shortVerizon bill, verizon.com−3.249612Likely legitimate
G32longprompt injection + Amazon refund+5.3100Likely scam
G33long"Your account has been tempor…" (truncated)−0.5381212Unrecognized (partial)
G34contactMom, preview hidden——95—Cannot assess (still alerts)
G35longunsaved climbing partner, prior substantive reply−1.5186238Likely legitimate
G36emailtoll follow-up after the user replied "Y"+6.0100Likely scam
G37shortCVS prescription ready−3.538512Likely legitimate
G38+44friend abroad−0.8312115Unrecognized
G39email"Hello, are you free this weekend?"−0.5381212Unrecognized (boundary)
G40longWells Fargo …secure-verify-login.com+5.5100Likely scam
G41longpolitical campaign text−1.5182112Unrecognized
G42longApple Support → AnyDesk+3.396Likely scam

What the corpus deliberately pins:


4. Scam-family coverage and sources

Family (US, 2024–2026)Codes that catch itEvidence
Fake package delivery (USPS/FedEx/UPS)PRETEXT_PACKAGE_PROBLEM, LOOKALIKE_DOMAIN, BRAND_DOMAIN_MISMATCH, INTL_CLAIMS_US_ENTITYFTC: #1 reported text scam of 2024; text-scam losses $470M in 2024 [1][2]
Unpaid toll (E-ZPass etc.)PRETEXT_TOLL, SMALL_FEE, PENALTY_THREAT, LOOKALIKE_DOMAINIC3 PSA I-041224-PSA: >2,000 complaints Mar–Apr 2024, with a template "outstanding toll $12.51… late fee $50" [3]. >60k toll complaints in 2024 [4]
Bank "did you authorize" / account lockedPRETEXT_BANK_ALERT, CALLBACK_NUMBER, CREDENTIAL_REQUEST, BRAND_FROM_LONG_CODEFTC top-5 (fake fraud alerts) [1]. Smishing Triad pivoted to bank kits in March 2025 [5]
Wrong number → pig-butcheringGENERIC_OPENER (0), WRONG_NUMBER_PIVOT, PLATFORM_SHIFT, INVESTMENT_PIVOTFTC top-5 [1]. IC3 2024: crypto investment fraud $5.8B, part of $9.3B in crypto-related losses [6]
Job / task scamsPRETEXT_JOB_TASK, PLATFORM_SHIFT, CRYPTO_PAYMENT_DEMANDFTC: task scams went from 0 reports (2020) to ~20k in H1 2024, ~40% of job-scam reports; >$220M job-scam losses in H1 2024 [7]
Reply-Y link activationREPLY_TO_ENABLE_LINK (T), THREAD_USER_REPLIED_TRIVIAL2025 campaigns telling users to "reply Y, exit and reopen" [8][9]
Phishing-as-a-service kits (Lighthouse, Darcula)Risky TLD list, LOOKALIKE_DOMAIN, email-gateway and international priorsUnit 42: >194k malicious domains since Jan 2024 [10]. Silent Push: 121+ countries [11]
OTP theft / account takeoverOTP_SHARE_REQUEST (T), OTP_DELIVERY, OTP_REDACTED_BY_OSAndroid 15 redacts OTP notifications from untrusted listeners [12]
Government (IRS, DMV, court), carrier, gift-card, family emergency, remote-accessPRETEXT_GOVT, PRETEXT_CARRIER, GIFT_CARD_PAYMENT, FAMILY_EMERGENCY_MONEY, REMOTE_ACCESS_APPThese are long-standing FTC consumer-alert patterns. Specific 2025 figures for these families were not re-verified in this review.

Not verified here, so treat as assumptions:


5. Calibration without a big dataset

5.1 Data sources

DatasetWhat it givesLimits
SmishTank (2024): 1,062 community-reported smishing messages with sender, brand, domain, VirusTotal and URL characterization [13]Real US-skewed lures with URLs and sender types. The best source for the LINK/PRETEXT/SENDER weights and the risky-TLD list.Positives only; no ham. Biased toward people who report.
Mendeley SMS Phishing Dataset (2022): 5,971 messages (4,844 ham / 489 spam / 638 smishing) [14]Has ham and smishing, so per-code likelihood ratios can be estimated directlyHam is old and non-US. Few modern business messages.
Mendeley balanced spam/smishing set for LLMs [15]More recent balanced labelsCheck provenance and language mix before use.
SMS Gateways smishing dataset (2023): 68k smishing messages crawled from public SMS gateways, confirmed via VirusTotal/APWG/GSB (cited in [16])Large-volume URL/TLD/brand statisticsSmishing only. Gateway-sourced, so skewed to OTP-gateway abuse.
Systematic review of SMS phishing (2026) [16]Pointers to further datasets
UCI SMS Spam Collection (2011)Ham baseline languageStale: pre-smishing-kit era, UK-centric. Use only for opener/ham language.
Hand-built benign US business set: build this, ~150 messagesReal reminders, carrier/bank/pharmacy/delivery short-code texts, vendor-domain clinic links, self-ID texts from strangersNo public dataset contains modern legitimate texts from unknown US senders. That is exactly the false-positive population this product must protect, so it has to be curated by hand (Byron's own phone plus volunteers, redacted).

5.2 Method

  1. Start from expert weights (this document).
  2. Estimate per-code likelihood ratios where a dataset has both classes: LR = (k_scam+α)/(n_scam+2α) ÷ (k_ham+α)/(n_ham+2α) with α=1 (Beta-smoothed).
  3. Shrink toward the expert weight: w = (n·log LR + m·w_expert)/(n+m) with m≈20 pseudo-observations. Codes seen rarely stay near the expert value.
  4. Estimate within families, not codes. Fit the family cap as the 90th percentile of the family's summed LR on positives. This directly measures the correlation problem from §1.2.
  5. Correct the prior for deployment base rate. Dataset class ratios are not deployment base rates. Set priors per sender kind from the user's own first 2–4 weeks in shadow mode: score silently and collect user labels. Remember the post-Google-filter shift: the deployed base rate is lower than any public dataset implies.
  6. The golden corpus is a regression gate, not a training set. Every weight change reruns tools/corpus.py. Any verdict change requires an explicit decision-log entry. Add a case for every user-reported misclassification.
  7. Report calibration honestly with a reliability diagram per sender kind once ≥200 labeled messages exist. Until then, show the UI percentages as banded words (low / elevated / high), not "73%". Two-digit precision would be false precision.

5.3 User feedback: adjusting safely

Feedback is itself an attack surface. A scammer can socially engineer "mark as safe" and then exploit the trust it earns.


Sources

  1. FTC Data Spotlight, "Top text scams of 2024" (Apr 2025): https://www.ftc.gov/news-events/data-visualizations/data-spotlight/2025/04/top-text-scams-2024
  2. FTC press release, "Text scams hit $470 million": https://oig.ftc.gov/news-events/news/press-releases/2025/04/new-ftc-data-show-top-text-message-scams-2024-overall-losses-text-scams-hit-470-million
  3. FBI IC3 PSA I-041224-PSA (toll smishing): https://ic3.gov/PSA/2024/PSA240412
  4. NBC, "Spam texts about unpaid tolls are on the rise" (60k+ IC3 complaints in 2024): https://www.nbcboston.com/news/national-international/spam-texts-about-unpaid-tolls-are-on-the-rise-heres-what-to-do-if-you-get-one/3659397/
  5. KrebsOnSecurity, "China-based SMS phishing triad pivots to banks" (Apr 2025): https://krebsonsecurity.com/2025/04/china-based-sms-phishing-triad-pivots-to-banks
  6. TRM Labs, key findings from the FBI 2024 IC3 report: https://www.trmlabs.com/resources/blog/a-record-breaking-year-for-cybercrime-key-findings-from-the-fbis-2024-ic3-report-k1c94
  7. FTC Data Spotlight on task scams (Dec 2024): https://www.ftc.gov/node/86951
  8. BleepingComputer, "Phishing texts trick Apple iMessage users into disabling protection": https://www.bleepingcomputer.com/news/security/phishing-texts-trick-apple-imessage-users-into-disabling-protection/
  9. Malwarebytes on the same campaign: https://www.malwarebytes.com/blog/news/2025/01/imessage-text-gets-recipient-to-disable-phishing-protection-so-they-can-be-phished
  10. The Hacker News, Smishing Triad / Unit 42 domains: https://thehackernews.com/2025/04/chinese-smishing-kit-behind-widespread.html
  11. Silent Push, "Smishing Triad": https://www.silentpush.com/blog/smishing-triad
  12. Android Authority, Android 15 sensitive notifications: https://www.androidauthority.com/android-15-sensitive-notifications-3416414
  13. Timko et al., "Smishing Dataset I: Phishing SMS Dataset from Smishtank.com" (2024): https://arxiv.org/html/2402.18430v2
  14. Mishra & Soni, SMS Phishing Dataset (Mendeley, 2022): https://data.mendeley.com/datasets/f45bkkt8pr
  15. Balanced spam/smishing dataset for LLMs (Mendeley): https://data.mendeley.com/datasets/vmg875v4xs
  16. "SMS Phishing Attacks and Defenses: A Systematic Review" (2026): https://arxiv.org/pdf/2604.11429

← Back to the proposal