musicleads
Find music businesses that provably have money, confirm we can reach them, and hand over a call list where every lead has a phone and an Instagram.
Software Signals
Every public fingerprint that proves a music business has real money, scored on strength and detectability.
musicleads/ folder is music industry only — musicians, producers, engineers, studios, labels, managers/agencies. No other verticals. Every signal, scraper, and target list here stays music-specific.Building block #1 of the musicleads system. The job of this doc: catalog every piece of software / subscription / service a music business buys — plus every public spend-behavior and output signal (ads, ticket prices, review counts, raised dollars, credits) — that implies real revenue, then separate the ones we can actually detect at scale from the ones we can't. Block #2 (the scraper) gets built off the "build-first" shortlist at the bottom.
The intuition: "If they pay for X, they probably make real money." Correct — but only for the right X.
The precise rule is not "pays for software → has money." It's:
That distinction is the whole game. Linktree fails the test — it's the flagship example but it's actually a weak signal. Broke hobbyists buy Linktree Pro constantly; there's no $90 tier (top is ~$24/mo), and the free tier is everywhere. It tells you almost nothing about revenue.
The strong signals are tools that are one of:
- Invite-only / vetted (you have to be making money to even get access) — AWAL, Stem, DISCO, a booking agency
- Enterprise-priced (nobody pays this "just in case") — Chartmetric Enterprise, Community.com SMS, Shopify Plus, Kajabi
- Only-makes-sense-at-scale (the tool is useless until you have real volume) — tour-management software, royalty-accounting software, a self-owned publishing entity
The lettered sections (6A–14C) extend the same rule beyond purchases: money leaves public exhaust — ad libraries, ticket prices, review velocity, raised dollars, credited work. Same bar applies: it only counts if it's pointless or impossible below a revenue line.
Everything below is scored against that rule.
The bio link is the single most-scrapable artifact a music business has. The tool they use, and especially the smart-link domain, leaks their tier. (Full domain→distributor map in §15 — that's the centerpiece.)
linktr.ee/… in biobeacons.ai/…bio.link/…, snipfeed.co/…komi.io/…stan.store/…ffm.to (company site feature.fm — NOT features.fm, unrelated)lnk.to (also bio.to, smarturl.it) — found.ee & li.sten.to are NOT Linkfire, see §15hypeddit.com, push.fmorcd.co, untd.io, vyd.co, empi.re, ada.lnk.to — see §15__NEXT_DATA__ JSON blob containing account.tier — live-confirmed on a real profile returning "paid2" plus the account-creation timestamp (§19). Fetch page → parse JSON → read tier. No logo heuristic needed (keep the footer-logo tell as fallback; map tier values to plan names before hard-coding). Upgrades ⚪ → 🟡 at best, but it's free precision — and it's exactly the "pays for Linktree" thesis, made scrapable.Who distributes an artist is a graded ladder. The bottom is pay-to-play (anyone). The top is invite-only (you're already making money). And it's detectable because the distributor stamps its name in public metadata.
hyperfollow.com) links; ℗ line "DistroKid"label: search) across 2+ artists, ℗ line showing DistroKiduntd.io/… linksempi.re/…; ℗ "EMPIRE"orcd.co/…; "Provided to YouTube by The Orchard"bfan.link; ℗ "Believe"ada.lnk.to/…; ℗ "ADA"vyd.co/…copyrights field / Apple ℗ linePaying to collect publishing money means there's publishing money to collect.
Everyone in this category screams money. The seats themselves live behind logins — but they leak through job descriptions: any org hiring names its stack ("proficiency with Chartmetric required"). Detection = JD text mining (§10A). Honest coverage note: this only catches orgs actively hiring — an enrichment pass, never a primary sweep.
songstats.com artist links in DJ bios (Pro tier not provable from link alone)(Spotify/Apple for Artists: free, not a signal — deliberately absent from the catalog.)
Free tiers are noise. But SMS platforms and big email lists are enterprise-priced because they scale with a real fanbase.
list-manage.com embedsstatic.klaviyo.com, klaviyo.js on sitelaylo.comcommunity.comIf they're running a store, there's money moving through it. Detectability here is excellent (tech fingerprints).
cdn.shopify.com, *.myshopify.com, powered-by: Shopify header (X-ShopId is legacy — gone on modern storefronts)*.bandcamp.comA music business paying to run Meta ads has a marketing budget and usually the sales to justify it — nobody sustains ad spend that isn't working. Better still, Meta runs a public transparency tool (the Ad Library), so this is one of the very few "spends real money" signals that's genuinely scrapable. It's also a pre-qualifier for our own upsell: an advertiser is already sold on paying for growth.
The signal is not "runs an ad" — a hobbyist boosts one post for $6 (the Linktree trap again). The signal is sustained, multi-creative, conversion-driven advertising.
eu_total_reach — a single estimated integer (NOT a bucketed range; spend/impression ranges stay political-only)connect.facebook.net/…/fbevents.js, fbq('track','Purchase')Spend-proxy tiers (how we bucket without a dollar figure):
- ⚪ Dabbler — 0–1 ad, short-lived → ignore.
- 🟡 Active — 2–4 ads, running ≥2 weeks → qualified lead.
- 🟢 Serious spender — 5+ concurrent creatives, running continuously, product/catalog CTAs, Pixel firing purchases → hot lead (proven, ongoing marketing budget).
How we detect it at scale:
- Meta Ad Library (public UI) —
facebook.com/ads/library, filter All ads + country, search by Page or keyword. Shows every active ad, its start date, and platforms (FB/IG/Messenger). Scrapable — apify has Facebook Ad Library actors (confirm the specific actor at build time). Remember: active-only, so duration = start date → now. - Ad Library API — official; full coverage for EU-served ads (DSA reach data) and political globally; commercial non-EU coverage is limited, so the UI/scrape is the US workhorse.
- Meta Pixel web fingerprint — crawl the site for
fbevents.js+fbq('init')+ standard events (Purchase,AddToCart,InitiateCheckout). Purchase events = a conversion advertiser with a real store. Same detection class as Shopify/Klaviyo (§6).
Killer stack: sustained Meta ads (§6A) + Shopify (§6) + a distributor smart-link (§15) = a music business that sells product, spends to acquire customers, and is signed to a real distributor. Top-decile lead.
Same logic as §6A, different transparency surfaces. The EU's DSA forced every major platform to run a public ad repository; coverage varies by region, but presence in ANY library = active ad budget.
adstransparency.google.com — search by advertiser/domain; shows verified advertiser, formats, last-shown dates. No $ for commercial (same caveat as §6A)linkedin.com/ad-library — search by company name(Snapchat US commercial ads and Spotify Ad Studio / SoundCloud Promote spend: no transparency surface → Appendix A.)
Build note: same tiering as §6A (dabbler / active / serious) — count creatives × duration across libraries. A business active in two+ ad libraries simultaneously (e.g. Meta + Google) is a confirmed serious spender, no further proof needed.
Tour-management software is useless until you have a crew, dates, and routing. Its presence = you tour = you make money.
Tour money is computable from fully public parts. This is the closest thing to reading their P&L off the street.
For the producer segment, a stocked, Pro-tier store is the signal — not a free profile.
beatstars.com/… store, Pro badgeairbit.com/…traktrain.com/…(Splice/LANDR subscriptions: invisible AND ⚪ weak — Appendix A.)
Service marketplaces publish both the rate and the completed-work count. reviews × listed price = a hard, computable revenue floor (reviews undercount orders, so it's conservative). This is the single best signal for the engineer/producer segment.
engineears.com profilesairgigs.comA producer running a course or coaching funnel is a producer making money — and the platforms are detectable.
kajabi.com/mykajabi.com fingerprint on sitestan.store, gumroad.com/…whop.com/…patreon.com/… + patron countEverywhere else we infer tiers. Crowdfunding pages print the actual number, publicly, forever.
qrates.com shows a suspension noticegraphtreon.comBooking, contracts, invoicing, payroll, accounting = a business with revenue to manage.
*.as.me, calendly.com/…, Square bookingaspmx.l.google.com (+alt1-4) OR single-record smtp.google.com (post-2023 default) — match bothNobody hires without revenue. Bonus: this is where all the "invisible" software from §4/§12 leaks — job descriptions name the stack ("experience with Chartmetric/DISCO/Curve required").
The studio segment's strongest public exhaust: a commercial space generates reviews, listings, and rates.
The tier of the DAW/plugin stack, and premium formats, separate hobby from commercial.
Retainers and pro catalog tools = agency or funded artist.
disco.ac share links in EPKs/sites/bios + JD mining (§10A)f.io / frame.io review links in posts/bios (sparse)(Highnote/Byta: private-by-design sends, zero public artifact — Appendix A.)
Not software — but the highest-trust public proof of professional income, all programmatically searchable. Doubles as the verification layer for self-claimed bios ("platinum producer" → check the DB).
riaa.com/gold-platinum searchtunefind.com song/show searchTaking an advance requires provable royalty income — the strongest possible money proof. Quiet deals are invisible (Appendix A), but two surfaces ARE public: deal announcements and auction listings.
royaltyexchange.com listing scrape (Sound Royalties deals stay quiet → Appendix A)Protecting a brand costs money and signals they expect it to be worth protecting.
tmsearch.uspto.gov (TESS retired Nov 2023)Everything here comes from one crawl of their domain + DNS. Cheap to run over an entire lead list.
js.stripe.com, buy.stripe.com links, PayPal/Square SDK/.well-known/apple-developer-merchantid-domain-association/products.json + /collections.json (often public)kl._domainkey, km/kt/ks; legacy kl1/kl2)klaviyo._domainkey)Not purchases — public statuses that either require a distributor/label deal or mark a crossed monetization threshold.
label:"Name" searchlabel: filterFiltering out no-money accounts is worth as much as finding money. Run this pass LAST over every candidate list.
.wixsite.com, sites.google.com, .godaddysites.com URLThis is the best signal we have because it is both publicly visible in the bio and reveals who distributes them (= their tier = their money). One scrape yields both "here's a music business" and "here's how big they are."
linktr.ee, beacons.ai, bio.linkGeneric link-in-bio · ⚪ none — anyonehyperfollow.comDistroKid · ⚪ entry-level distroffm.toFeature.fm · 🟡 marketing budget (match ffm.to only; features.fm is unrelated)lnk.to (also bio.to)Linkfire · 🟡 serious artist / small labelfound.eeDowntown Music (CD Baby ecosystem) — NOT Linkfire · 🟡 Downtown-distributedli.sten.toListenTo — independent tool, NOT Linkfire · ⚪ generic smart-linkuntd.ioUnitedMasters · 🟡 SELECT tier + deal flowbfan.linkBelieve · 🟢 large indie machineempi.reEMPIRE · 🟢 real dealsorcd.coThe Orchard (Sony) · 🟢 vetted premium indieada.lnk.toADA (Warner) · 🟢 vetted, Warner servicesvyd.coVydia · 🟢 enterprisesmarturl.itLegacy (often major/Universal) · 🟢 usually major-adjacentfanlink.toToneDen · 🟡 marketing-tool user (free tier exists — verify)song.link / odesli.coOdesli · ⚪ none — free toolorcd.co→Orchard, untd.io→UnitedMasters, vyd.co→Vydia, empi.re→EMPIRE (links on music.empi.re), bfan.link→Believe, smarturl.it→Linkfire-acquired, song.link→Odesli — all confirmed by resolving live links. Corrected: found.ee=Downtown (not Linkfire), li.sten.to=ListenTo (not Linkfire), features.fm≠Feature.fm. Nuance: a branded domain identifies the distributor/servicer of that release (strong proxy), not cleanly the artist's tier or current deal — and it only helps for the subset who use branded links; its absence proves nothing (most use neutral tools like linktr.ee).The strongest money signals hide behind logins — but most of them leak anyway. The v1 draft wrote off Chartmetric seats, Kobalt admin, QuickBooks payroll, and DISCO as "invisible." The rigor pass found programmatic leak paths for nearly all of them:
- PRO repertory admin records — Songtrust/Kobalt/Sentric administration is printed on public ASCAP/BMI work records
- JD text mining — every pro tool a hiring org uses gets named in its job posts (§10A)
- Public share links —
disco.ac,f.ioin EPKs/bios - Gear-list pages — studios publish their own software/equipment lists (§11)
- Auction listings — Royalty Exchange prints actual last-12-month earnings (§13)
The catalog now contains only programmatically detectable signals: 🟢 = direct fingerprint, 🟡 = multi-step or sparse leak path. The handful with genuinely zero public artifact live in Appendix A, banned from the pipeline so no build-hour is ever spent on them.
At-scale discipline: 🟢 direct fingerprints are primary sweeps over the whole candidate pool; 🟡 leak paths are enrichment passes on survivors only (see the funnel corollary at the top).
Grouped by method, because Block #2 is built method-by-method, not signal-by-signal.
- Bio-link harvesting (our primary engine). Instagram bios contain the link-in-bio URL. We have the apify Instagram scraper — bio+
externalUrlcome back single-stage from profile/user SEARCH or profile URLs withresultsType=details(hashtag/location are 2-step: collect handles → fetch profiles). Extract the domain, classify against §15. This is the connective tissue of the whole system: scrape IG → extract bio domain → classify tier → keep the 🟢s. - Web tech fingerprinting (BuiltWith/Wappalyzer-style, or roll our own): crawl their site, detect Shopify / Klaviyo / Kajabi / HubSpot / booking tools from HTML + headers.
- Streaming metadata: Spotify API
copyrightsfield and the "Provided to YouTube by …" line reveal distributor/label → tier (§2, §15). - DNS/MX lookup: custom domain + Google Workspace MX = real business (§10).
- Public databases: USPTO Trademark Search (
tmsearch.uspto.gov; TESS retired), state registries / OpenCorporates (LLCs; bulk = paid API), ASCAP/BMI/SESAC repertory (self-owned publishing), SoundExchange (§3, §14). - Directory scraping: SoundBetter (priced pros), BeatStars/Airbit (stocked producer stores), Bandsintown (touring), Discogs/AllMusic (catalog depth).
- Press/agency detection: reputable press hits + WME/CAA/UTA mentions in bio/press (§7, §12).
- Ad-library sweep — all platforms: Meta Ad Library (apify actor) + Google Ads Transparency Center + TikTok CCL (EU) + LinkedIn Ad Library, by page/advertiser/keyword, for active-ad count & duration; plus Meta Pixel + Purchase-event crawl on their site. Spend proxy, not exact $ — see §6A/§6B.
- Job-board scrape + headcount: Indeed / LinkedIn Jobs / EntertainmentCareers for live postings; LinkedIn company pages for employee count. JDs also out the invisible §4/§12 software (§10A).
- Maps & booking marketplaces: Google Maps reviews scrape (dates → review velocity) for studios; Studiotime/Peerspace/Giggster rates × reviews (§10B).
- Crowdfunding & fan-funding pages: Kickstarter/Qrates exact raised $; Patreon tiers × patrons (+ Graphtreon history); Bandcamp collector walls (§9A); Royalty Exchange auction listings — actual earnings (§13).
- Track-record databases: RIAA cert search, Muso.AI/Jaxsta/Discogs credits, Tunefind syncs, awards DBs, SESAC/GMR + ASCAP/BMI repertory including admin-entity parse (Songtrust/Kobalt/Sentric show as administrators on work records) (§12A, §3).
- Payment & app forensics: Stripe/
buy.stripe.com/Apple Pay well-known file, Shopifyproducts.jsondepth, RDAP domain age, app-store search, podcast RSS hosts, studio gear-list keyword parse (§14A, §11). - Platform badges & live math: YouTube OAC/Vevo/Join, Spotify
label:release cadence, IG paid-partnership labels (§14B); tour gross = ticket price × venue capacity × dates (§7A). - Bio-claim NLP + verification loop: extract claims ("platinum", "Grammy-nominated", "as seen in Billboard") from scraped bios → verify against §12A databases. Claim alone = 🟡; DB-verified = 🟢.
- Negative-signal filter pass (§14C): runs LAST over every candidate list — kills free-tier subdomains, dormant catalogs, dead funnels, bought audiences.
The only signals that go into Block #2's first scraper. Ranked by ROI = (money strength × detectability × how much it narrows to our customer).
orcd.co, ada.lnk.to, vyd.co, empi.re, untd.io)copyrightsproducts.json (§6, §14A)label: release cadence (label segment)tmsearch.uspto.gov) + press scrape (§14, §12)Then the §14C negative-signal pass runs over everything — free-tier subdomains, dormant catalogs, dead funnels, bought audiences get culled before a human ever sees the list.
Not in the top-16 but still in the catalog as enrichment passes (programmatic, just multi-step/sparse — run on funnel survivors, not the whole pool): PRO-repertory admin records (Songtrust/Kobalt/Sentric), JD mining for pro stacks (Chartmetric, DISCO, QuickBooks, Master Tour), disco.ac/f.io share-link crawls, gear-list parsing, Royalty Exchange auction earnings, announced advances. The truly-undetectable residue lives in Appendix A and stays out of the pipeline entirely.
And §19 is the calibration layer: the Burn County archetype (money exhaust × presentation debt) converts these raw signals into the profile of the one lead that actually closed.
Why this section exists. Ty Weathers (Burn County / Sand City Studio, North Myrtle Beach SC) went from IG-discovered lead → paying Studio Raine client → the source of the RIAA-Gold and Billboard credentials now displayed on studioraine.art. He is the one fully-closed loop we have: every signal below was ON him at lead time, so this is the empirically-validated fingerprint — not theory. Everything here was live-verified July 24, 2026 (raw fetches of linktr.ee/burncounty, all four of his domains, DNS/MX lookups, revenue-link status checks) and cross-checked against the July 22 deep-research audits in research/.
leadconnectorhq ×90–160 + msgsndr assets in HTML of sandcitystudio.com, burncountymusic.com, burncountymedia.com, tyweathers.com; www CNAME → sites.ludicrous.cloud on all four; A → 162.159.140.166pipelinepro.co reseller instance)link.pipelinepro.co/widget/form/… + /widget/booking/…/widget/(form|booking|survey)/{id} in any bio link = GHL user, regardless of white-label domain — no site crawl needed__NEXT_DATA__ → account.tier: "paid2", account since Jan 2022skool.com/@ty-weathers-9557paypal.me/burncounty tile; stripeAccountId config in page JSThe two-axis insight this client proves: the lead is not "has money" alone — it's "has money AND has a visible, provable presentation gap." Ty was Top 1% on paper and invisible online. Every gap below is programmatically detectable, which means the outreach hook itself is scrapable (we opened with his own broken funnel).
DB_hit(RIAA/Muso/Wikipedia) == TRUE AND site_text.contains("RIAA|gold|Billboard|platinum") == FALSE → "Top 1% on paper, invisible online." Both halves already in the catalog (§12A × site crawl) — the JOIN is the signal/home-7753-8909, /portfolio-3446-966539, /link-in-bio-6468-5408/[a-z-]+-\d{3,4}-\d{4,6}$ = cloned funnel pages, no designed siteburncounty.com printed on the EPK was a dead domaincognitoforms.com/BradCox2/… under a personal name, "Powered by Cognito Forms"cognitoforms.com|typeform.com|jotform.com|forms.gle for a core paid serviceeforward*.registrar-servers.com (Namecheap forwarding) on 2 domains, no MX at all on burncountymusic.com — while printing info@burncounty.com everywherefbevents.js = 0 across all four sites — an agency owner running zero retargetingfbevents.js/gtag = unsophisticated funnel (flip side of §6A)user_ratings_total == 0 + business age (inverse of §10B velocity)IG_followers < 500 AND DB_hit == TRUE (variant of #1)GoHighLevel ($97–$497/mo) is the highest-value single software signal this client validated, and it's detectable three ways:
sites.ludicrous.cloud (+ legacy GHL targets — enumerate at build). Filter the domain list by music terms (studio, records, recording, mastering, beats, entertainment…) in domain/title/homepageleadconnectorhq|msgsndr) on hits/widget/(form|booking|survey)/{id} = GHL behind ANY white-label domain (that's how pipelinepro.co outed itself)- GHL Stripe false positive: every GHL page ships
stripeAccountId/stripePublishableKeyplumbing (18 hits on all four Ty sites, identical count = platform boilerplate). Do NOT count §14A "live payment rails" on a GHL site without a populated account value or a real order form. - Lead-time vs post-fix state: detectors describe the LEAD-TIME snapshot. Ty's
burncounty.comwas dead at lead time — today it serves John's rebuild (title: "Sand City Studio | Gold-Record Recording Studio…"). Post-engagement, the archetype's problem signals flip OFF — which is both the proof the play works and a reminder that stale scrapes = stale hooks. Re-verify hooks within days of outreach. - Linktree tier mapping:
paid2observed live; map the full value set (free/paid1/paid2/paid3 → plan names) before hard-coding thresholds. - The booking 404 is still live (July 24) — broken revenue links persist for months in this segment. The hook inventory doesn't go stale fast.
These cannot be programmatically detected by any known path. Listed only so nobody re-adds them or burns a build-hour trying. If a leak path ever appears (a transparency law, a public library, a share-link format), they graduate back into the catalog.
- Target segment first? The scraper targets one segment at a time. Artists, producers, or studios/engineers as the beachhead? (Changes which signals lead.)
- Who's the buyer at the end — are these leads for Studio Raine's services, or for Sand City Studios? Determines what "qualified" means.
- Verification pass — want me to run the live §15 domain→distributor verification next, since it's the hard blocker on the #1 signal?
Notes
Log every callback, every lead, every follow-up. Saved on this device.
Contactability
Turning money-qualified candidates into a callable, textable list. Phone and Instagram, every row.
musicleads/ folder is music industry only — musicians, producers, engineers, studios, labels, managers/agencies. No other verticals.Building block #2 of the musicleads system. Block #1 (softwaresignals.md) answers "who has money and needs us?" This block answers the very next question: "can John actually reach them — by phone and Instagram — easily, at 100 leads/day?" A lead John can't call is worth zero, no matter how gold-plated the money signals. So contactability is not a nice-to-have downstream step; it's a hard gate that sits between the money-qualified candidate pool and the final call list.
Every row that reaches John's call list must clear both:
The phone isn't just a contact detail — a business that publishes a phone number for calls is self-selecting as easy to sell to. If they take calls publicly, they take sales calls. That makes "how the phone is exposed" a qualifying signal, not just a lookup:
Two routing facts that matter for John's outreach:
- IG business/creator "Call" button + Google Business phone = a business saying "call me." Prioritize these.
- Mobile numbers are textable; landlines are call-only. Line-type detection (optional enrichment, §7) lets us tag each row Call vs Call+Text so John texts only what's textable.
Across this niche the contact chain is nearly always the same, and every hop is machine-followable:
Instagram profile ──▶ bio link (Linktree / link-in-bio / direct URL) ──▶ website ──▶ PHONE listed for calls
│ │
└── handle (we need this) └── broken? (serveability) + phone (contactability)Both things John needs — the IG handle and the phone — fall out of walking this one chain. The system below is just that walk, hardened, deduped, validated, and run in batch.
Input is the money-qualified candidate pool from Block #1 (the §17 scraping bridge output / §19 Burn County archetype). Output is a complete, deduped call list. Stages run as a funnel — cheap steps on everyone, expensive steps only on survivors (same discipline as softwaresignals.md).
[Block #1 output: money-qualified candidates]
│ each has SOME identifier: IG handle OR domain OR Maps place OR Spotify id
▼
STAGE 1 — IDENTITY RESOLUTION → normalize every candidate to {IG handle, website domain}
▼
STAGE 2 — PHONE WATERFALL → try sources in order; collect all; cross-validate (§5.3)
▼
STAGE 3 — INSTAGRAM CONFIRM → handle is real, public, active, ideally professional (§5.4)
▼
STAGE 4 — SERVEABILITY GATE → website broken/weak/absent? (else DROP) (§5.5)
▼
STAGE 5 — CONTACTABILITY GATE → has phone AND IG? (else DROP) (§5.6)
▼
STAGE 6 — VALIDATE & ROUTE → E.164 normalize, line-type, call vs text, dedupe (§7)
▼
[Final call list: phone + IG + hooks, 100/day]Every candidate must be reduced to two anchors: an Instagram handle and a website domain. Where each starts depends on how it entered from Block #1:
externalUrlsocialMedias/website social-icon scrapeinstagram.com/ hrefJoin key for dedup: normalized domain first, else IG handle. The same operator often appears via IG and Maps and domain (Burn County was on all three) — collapse to one record here so John never sees a dupe.
IG-handle and domain are effectively free to carry. The paid enrichment (profile scrape, Maps place, site crawl) happens in Stage 2's waterfall and is ordered cheapest-first so most leads resolve before the expensive calls fire.
Try sources top-to-bottom; first valid hit wins, but capture every source that returns a number so we can cross-validate (agreement across two sources = high confidence). All are batchable at 100+/day.
logical_scrapers/instagram-profile-scraper → allPhoneNumbers[] (also businessPhone, businessContactMethod, hasContacts, isProfessionalAccount)compass/crawler-google-places → phone / phoneUnformatted (or lukaskrivka/google-maps-with-contact-details)tel: hrefs, JSON-LD telephone, footer/contact regex — or batch purple_beep_boop/bulk-website-contact-scraper → phone__NEXT_DATA__ + "Call"/"Text" tiles (tel:/sms:)Coverage reality (honest): no single source covers everyone. IG-phone only exists for professional accounts with a contact button; Maps-phone only for businesses with a listing; website-phone only if they list one. That's why it's a waterfall — the union of sources is what gets us to near-100% on the archetype. For the Burn County profile specifically, sources 1, 3, and 4 all hit (IG business acct, GHL site tel: button, Linktree). Redundancy is the norm for this niche, not the exception.
apify/instagram-scraper does NOT return a phone field in its output — confirmed against its schema — so phone requires a contact-capable actor like logical_scrapers/instagram-profile-scraper (field allPhoneNumbers). (b) businessPhone came back null in that actor's sample runs while allPhoneNumbers was populated — key off allPhoneNumbers, treat businessPhone as a bonus. (c) Confirm exact field names + current pricing on a 10-lead test run before scaling.IG is required output, and in most paths (§5.1) we already have the handle. Confirm it's usable with one profile scrape (reused from Stage 2 #1):
- Exists + public (not private, not deleted) —
isPrivate == false, scrape succeeds. - Active — has a recent post (reuse the §14C dormancy check: last post < 6 months).
- Ideally professional —
isProfessionalAccount == true→ unlocks the §5.3-#1 phone AND is a mild money/seriousness signal. - Capture
followers,bio,namefor the output row and the outreach hook.
John sells website creation, so a great site disqualifies. Score website brokenness; pass = broken/weak/absent. This reuses the §19.2 problem-signal detectors, as a machine rubric:
leadconnectorhq/msgsndr, Wix, Weebly, GoDaddy Sites), auto-slug URLs (/home-7753-8909), broken/404 internal links, no <meta viewport> (not mobile-ready), no HTTPS, single-page "call for quote," dated © 20xx, stock-only imageryThe non-negotiable filter, run last before validation: DROP any record missing phone OR IG. Only complete {phone, IG} rows survive. This is what guarantees John's promise to himself — every single lead on the list has both.
All confirmed live against the Apify store (July 24, 2026), with real fields/prices. We already have the apify MCP wired in.
logical_scrapers/instagram-profile-scraperallPhoneNumbers[], allEmails[], bioLinks[], websiteLinks[], socialLinks[], isProfessionalAccount, hasContacts, businessCategory, followers, bioapify/instagram-scraperbiography, externalUrl, isBusinessAccount, businessCategoryName, followersCount, verifiedlukaskrivka/google-maps-with-contact-detailscompass/crawler-google-placesphone, phoneUnformatted, website, title, reviews, scrapeContacts add-onpurple_beep_boop/bulk-website-contact-scraperBad numbers waste John's call time, so every surviving number is cleaned:
- Normalize to E.164 (
+1XXXXXXXXXX); strip formatting. - Sanity-check — valid NANP area code, 10 digits, not
555/000, not an obvious platform/toll number, not a fax label. - Cross-validate — if two sources returned the same number → 🟢 high confidence; if they disagree → keep both, flag
phone_conflictfor John to eyeball. - De-dupe on E.164 phone AND on the Stage-1 join key (one operator = one row).
- Line-type (optional) → tag
Call(landline) vsCall+Text(mobile) so John only texts textable numbers. - Guard the GHL-Stripe-style false positives (per §19.5): a
tel:that resolves to a platform default, not the business, gets dropped.
The studio segment is the beachhead for contactability: Maps gives phone+site+IG together, and local studios have the most broken websites — both gates clear at once.
Not every money-qualified candidate clears both gates, so we over-feed the funnel. Rough, honest math (tune with real run data):
- Assume ~40–55% of Block #1 archetype candidates end up with phone AND IG AND a broken/absent site.
- To net 100/day, feed ~200–250 candidates/day into Stage 1.
- Per-candidate enrichment cost (IG profile + maybe Maps + maybe site contact): ~$0.01–0.03.
- → ~$3–7/day in apify to produce 100 fully-contactable, serveable leads. Trivial versus one closed website deal.
The bottleneck is Block #1 candidate supply, not contactability tooling — the actors easily handle thousands/day. So Block #3 (the actual scraper) should target ≥250 archetype candidates/day to keep this funnel full.
Every row is complete by construction (both gates passed). This is the deliverable format:
business_nameoperator_namephone_e164+18435803057phone_routeCall+Text (mobile) / Call (landline)phone_sourceIG contact button | GBP | website tel:contact_tierinstagram@sandcitystudioig_followerswebsitewebsite_verdictmoney_signalsoutreach_hooksegmentThe last two columns are why this system compounds: the outreach hook is carried straight from Block #1's problem signals (the 404, the missing gold record) — so John opens every call already knowing the lead's exact wound.
The system John asked me to create, in one paragraph: take Block #1's money-qualified candidates, resolve each to an Instagram handle + a website domain, then run a phone-acquisition waterfall (IG professional-account phone → Google Business phone → website tel:/contact scrape → Linktree tile → Facebook → WHOIS), keeping the first valid number and cross-validating any others. Confirm the IG is live and public. Gate on serveability (website must be broken/weak/absent — John sells websites) and then on contactability (must have phone AND IG). Validate/normalize the number, tag it Call vs Text, dedupe, and emit a row carrying the phone, the IG, the money signals, and the outreach hook. It runs as a cheap-first funnel on ~250 candidates/day to net 100 complete leads for ~$3–7/day, using the verified apify actors in §6 (logical_scrapers/instagram-profile-scraper for IG phone, lukaskrivka/google-maps-with-contact-details for studios, purple_beep_boop/bulk-website-contact-scraper for the site hop). The design deliberately fuses the two gates: the broken website that qualifies them for John's #1 service is the same artifact that hands us their phone number.
- Beachhead segment? Studios clear both gates most easily (Maps = phone+site+IG in one call, and their sites are the most broken). Recommend starting there — confirm?
- Line-type routing — worth wiring Twilio Lookup (~$0.005/#) so every row is tagged Call vs Text, or is call-only fine for v1?
- "No website at all" leads — keep them (they need a site from scratch, §5.5) or does John only want visibly-broken-site leads so the pitch has a before/after? Recommend keeping both, tagged.
- Phone-confidence floor — ship 🟡 callable (website-listed) rows, or hold the first list to 🟢 hot-to-call only (IG/GBP contact button) for the highest answer rate?