Automations · connected capabilities

AI automation tools. What did each contribute?

How Make, Apify, Firecrawl and Exa helped turn prospect searches into business research. Explore the AI automation workflows for finding businesses, assessing websites and analysing social profiles—with the outputs, costs and practical lessons from each tool.

Follow the workflows →
Specialist tools. Evidence brought together.
Find prospects→Capture & qualify websites→Collect profiles & posts→Compare businesses
Apify9,806items across five principal tasks

Company profiles and posts for Facebook, LinkedIn and TikTok.

Explore Apify ↗
16,345recorded operations

Instagram profile and post material connected to business records.

Explore Make ↗
Firecrawl559 / 607difficult-site domains captured

A 92.1% capture result in the named August qualification batch.

Explore Firecrawl ↗
Exa687missing-site searches

236.2 seconds and $4.809 in the named discovery batch.

Explore Exa ↗
1,478returned result records

1,386 records with resolved accounts; 1,298 with a latest-post date. 3,049 requests across 1,622 submitted candidates.

Explore X ↗
YouTube766returned result records

616 records with resolved channels; 582 with a latest-video date. 820 submissions used 1,956 quota units.

Explore YouTube ↗
Website evidence6,131distinct website captures

Distinct HTML content hashes recorded across the combined evidence register, covering 6,562 business records. Includes multiple capture routes; the August direct-retrieval campaign is detailed below.

Explore website evidence ↗

Inside the collection.

Open a tool to see its workflow, returned fields, recorded results and costs.

Apify

Apify supplied the detail behind the social profiles—not just a list of accounts

Five principal collection tasks returned 9,806 items in the retained September receipts. They supplied company descriptions, website links, audiences, opening hours, service areas, employee information and post-level activity. The practical value was joining those details to the website assessment for the same business.

Apify’s reported programme cost was $55.07. The retained September consumption ledger covers $31.48 of Apify work; the five principal tasks below account for $30.57 within that ledger. These are different scopes, not competing totals.

Facebook took most of the principal-task spend

Collection task Actor used Submitted Returned items Recorded cost
Facebook page details apify/facebook-pages-scraper 3,579 3,277 $15.09
Facebook posts apify/facebook-posts-scraper 3,610 3,580 $12.59
LinkedIn company details harvestapi/linkedin-company 1,593 1,495 $0.87
LinkedIn company posts harvestapi/linkedin-company-posts 1,593 1,055 $1.37
TikTok profiles clockworks/tiktok-profile-scraper 400 399 $0.65

Facebook’s two tasks account for $27.68—about 91% of the spend in this five-task comparison. That makes their output quality the first place to look when judging value from this part of the project. The profile and post tasks also answer different questions: who the business is and how it describes itself, versus what it has been publishing.

Facebook details: contact routes, service information and audience

Captured field Records with a value / returned records What it contributes
Page identity 2,849/3,277 Links details to the correct page
Page name 2,849/3,277 Matches the business identity
Categories 2,800/3,277 Describes the work offered
Intro 2,708/3,277 Short explanation of the business
Website 2,730/3,277 A route from social profile to website
Email 2,486/3,277 A published contact route
Telephone 2,456/3,277 A published contact route
Followers 2,803/3,277 Audience benchmark
Page likes 2,803/3,277 Separate from follower count
Following 2,787/3,277 Profile context
Opening hours 1,903/3,277 When customers can use the service
Address 2,208/3,277 Location evidence
Service area 797/3,277 Geographic scope
Profile image 2,803/3,277 Profile presentation
Cover image 2,780/3,277 Profile presentation
Page creation date 2,800/3,277 Age of the account, not the business
Other social links 664/3,277 Connects the wider profile
WhatsApp number 322/3,277 Alternative enquiry route
Services 197/3,277 More specific service wording

Facebook supplied website values in 2,730 records, email values in 2,486 and opening-hours values in 1,903. This is substantially more useful to business research than a recommendation percentage that barely varies. These fields describe whether the profile offers a next step and explains when and where the business operates.

The retained output also includes recommendation and rating fields, page-advertising state, price-range information and further linked-platform fields. They remain available in the field catalogue; rating ceilings are not promoted into business-performance headlines. Account details are assessed after matching them to the business.

Facebook posts: what was being published and the response it received

Captured field Records with a value / returned records What it contributes
Post identity 3,145/3,580 Connects observations of the same post
Post date 3,145/3,580 Posting recency
Post text 2,681/3,580 Services, offers and projects described
Likes 3,145/3,580 Response to the collected post
Shares 3,145/3,580 Resharing of the collected post
Comments 557/3,580 Reported discussion count
Media 2,882/3,580 Photographs or other attached material
Outbound link 1,025/3,580 Destination offered to readers
Video views 535/3,580 Audience for video material where returned
Shared-post origin 132/3,580 Distinguishes original and reshared material

Post text was available in 2,681 records and media in 2,882. Together with the post date, this material gives the analyst the substance behind an activity figure: project photographs, service descriptions, offers and other business updates. The shared-post reference also allows a reshare to be distinguished from material the business published itself.

For a contractor, the valuable comparison is whether current work is visible and whether a post leads to useful information or an enquiry. Likes and views add context to individual posts. They are not a substitute for that assessment, and a latest-post sample does not measure the performance of an entire campaign.

LinkedIn details: business context that a follower count cannot supply

Captured field Records with a value / returned records What it contributes
Company identity 1,495/1,495 Matches the company page
Company name 1,495/1,495 Identity evidence
Tagline 1,240/1,495 Short positioning statement
Website 1,480/1,495 Company website reference
Phone 944/1,495 Contact route
Description 1,478/1,495 Services and specialism
Specialities 1,097/1,495 More precise activity labels
Industry 1,495/1,495 Category context
Locations 1,495/1,495 Geographic context
Employees 1,495/1,495 Company-size context
Employee-range lower bound 1,451/1,495 Reported size band
Employee-range upper bound 1,432/1,495 Reported size band
Followers 1,495/1,495 Audience size
Founded year 1,192/1,495 Reported company age
Company type 1,438/1,495 Organisation context
Logo 1,434/1,495 Profile completeness
Cover image 1,321/1,495 Profile presentation
Page type 1,495/1,495 Account classification
Verified-page flag 1,495/1,495 Platform-reported state
Call-to-action URL 1,478/1,495 Destination offered by the page

LinkedIn returned a website in 1,480 of 1,495 company-detail records and a description in 1,478. Founded year was available in 1,192 returned records. These are useful inputs to company-age and service comparisons once the company identity is accepted; they are not a reason to assume every record in the business master has a reliable founding date.

Employee count and company-type fields add a different perspective from consumer-facing platforms. A business with a small social audience may still present a substantial organisation through its company description, specialities and workforce information. That is why the study retains the detail rather than reducing LinkedIn to “account found”.

LinkedIn posts: company information and publishing activity are separate outputs

Captured field Records with a value / returned records What it contributes
Post date 1,055/1,055 Recency
Content 1,045/1,055 Message and service wording
Author type 1,055/1,055 Company/person context
Likes 1,055/1,055 Response to the post
Comments 1,055/1,055 Discussion
Shares 1,055/1,055 Redistribution
Images 1,055/1,055 Visual material
Article title 92/1,055 Linked editorial content
Video URL 113/1,055 Video format
Document title 34/1,055 Downloadable or document content
Reposted-content identity 53/1,055 Identifies a reshare

The post collection includes text in 1,045 records, video links in 113 and document titles in 34. Those formats help explain how businesses demonstrate expertise: short updates, longer articles, video and documents are different forms of presentation. An analyst can connect the material to the company’s stated specialities and website services, rather than guessing its marketing strategy from follower count.

TikTok: audience, video activity and the route beyond the platform

Captured field Records with a value / returned records What it contributes
Profile identity 367/399 Account matching
Biography 344/399 Service description
Bio link 93/399 Destination outside TikTok
Followers 367/399 Audience size
Following 367/399 Profile context
Profile likes 367/399 Accumulated platform response
Video count 367/399 Content volume
Verified flag 367/399 Platform state
Private-account flag 367/399 Access context
Latest video date 344/399 Recency
Video text 335/399 What the video describes
Video URL 344/399 Source reference
Video duration 344/399 Content format
Views 344/399 Video audience
Likes 344/399 Video response
Comments 344/399 Discussion
Shares 344/399 Redistribution
Saves 344/399 Recorded saves
Hashtags 344/399 Content labels
Pinned flag 344/399 Distinguishes pinned material
Sponsored flag 344/399 Platform-reported promotion context

The TikTok task returned 367 profile identities and 344 dated video records, while a bio link was populated in 93 records. That distinction matters to a business using video to attract attention: the profile and its content are one part of the journey; a destination for the interested viewer is another. At business level, the retained bio link can be checked against the website and enquiry routes already collected.

What we derived after collection

The research added eight completeness checks where the relevant field exists: phone, website, about text, logo, biography, email, reported hours and categories. It also classified ownership, retained collection outcomes and calculated posting recency. These checks make the raw fields comparable without pretending every platform has the same schema.

A business-facing table therefore shows the available accounts, the measured audience and dated activity, with a short platform note. The electricians edition demonstrates that approach. The practitioner’s output is the field catalogue and the collection results; the owner’s output is the comparison and a practical recommendation.

Download the aggregate field catalogue. It lists populated field counts, not business identities or profile values.

How collection outcomes were handled

Task Returned records Records carrying profile/post identity Explicit error records
Facebook details 3,277 2,849 428
Facebook posts 3,580 3,145 435
LinkedIn details 1,495 1,495 0
LinkedIn posts 1,055 1,055 0
TikTok profiles 399 367 32

Collection errors were retained as explicit outcomes alongside the usable profile and post records. They formed part of the operational account of the work: submissions, material returned, errors encountered and the evidence available for qualification. They were not treated as business attributes or used to erase the useful fields collected elsewhere.

For a practitioner, this makes the process reviewable. Facebook details delivered 2,849 page identities and Facebook posts delivered 3,145 post identities. LinkedIn supplied company detail and publishing evidence through separate tasks. Those outputs then supported the business-level ownership checks, audience comparisons, service analysis and posting timelines.

Sources and reading the figures

These figures come from retained project receipts and outputs. Field counts describe populated values in returned records, including repeated observations; they are not counts of unique businesses. A returned error item is reported separately from a returned profile. Blank fields are not converted into zero. Completeness flags, ownership decisions and qualification scores are derived after collection, not supplied as scores by the scraper. No additional collection was performed for this edition.

Source material: September provider-consumption receipts, platform raw-record exports in the combined research database, and the verified August qualification closeout where explicitly stated. Programme-wide reported costs and individual collection-window costs are kept separate. Identifiable source records remain private.

Make

Make connected Instagram profiles to their posts

The retained social receipts record 16,345 Make operations against 2,991 submitted candidates. The Instagram export contains 2,434 profile records, 7,205 post records and 414 error records.

Make’s job in this project was the Instagram collection workflow. It brought profile and post material into the research so that audience, biography, website and current activity could be joined to the same business. A single Instagram account could produce several post records; an operation was a workflow step, not a business or a post.

What the workflow returned

Retained output Records How it was used
Profile records 2,434 Audience, profile volume and identity
Post records 7,205 Dates, captions and response
Error records 414 Separate unsuccessful collection outcomes
Total Instagram export 10,053 Joined profile/post/error record set

The raw export mixes these record types. A row with a post caption is not another company profile. Keeping them separate allows a practitioner to compare the cost of collection with both the account detail and the material it delivered.

Fields available for the research

Captured field Records with a value / returned records What it contributes
Username 10,053/10,053 Connects profile and post output
Profile name 2,812/10,053 Identity matching
Followers 2,434/10,053 Audience
Posts published 2,434/10,053 Profile content volume
Following 2,434/10,053 Profile context
Website 2,657/10,053 Destination
Biography 2,762/10,053 Services and positioning
Post type 7,205/10,053 Content format
Post date 7,184/10,053 Recency
Caption 6,900/10,053 Project and service wording
Post URL 7,184/10,053 Source reference
Post likes 6,848/10,053 Recorded response
Post comments 7,184/10,053 Recorded discussion

The denominator here is the combined 10,053-row Instagram export. Profile fields and post fields belong to different record types; a field absent from a post row is not a missing field on the company profile.

The export has dated post values in 7,184 rows and captions in 6,900. Those fields provide the material behind a current-work comparison. Follower count describes the audience already accumulated; captions and dates show the work being presented to it.

In the electrician study, Instagram was available for 28 businesses, but only five had a post within 30 days. The automation’s commercial output is that distinction and the underlying project material. It allows a conversation about keeping an existing profile useful, rather than merely recommending Instagram because it is popular.

How the pieces fit together

The retained workflow submitted account candidates, collected profile and post outputs, saved them with the business reference and then normalised the fields for the social assessment. Ownership and collection outcome were handled before business-level comparisons. This let profile information, latest activity and website evidence appear together in the same business report.

For a practitioner, the reusable pattern is to preserve that join through the automation: business reference → account → profile fields → dated posts → report measures. Make’s operation count measures workflow consumption. The profile and post counts describe the evidence delivered. The report’s business counts describe the final analytical population.

The reported programme cost for Make was $18.82. That cost and the 16,345-operation receipt total have different accounting scopes, so this edition does not manufacture a cost per successful business from their division.

Aggregate field catalogue · Electrician audience and activity comparison

Sources and reading the figures

These figures come from retained project receipts and outputs. Field counts describe populated values in returned records, including repeated observations; they are not counts of unique businesses. A returned error item is reported separately from a returned profile. Blank fields are not converted into zero. Completeness flags, ownership decisions and qualification scores are derived after collection, not supplied as scores by the scraper. No additional collection was performed for this edition.

Source material: September provider-consumption receipts, platform raw-record exports in the combined research database, and the verified August qualification closeout where explicitly stated. Programme-wide reported costs and individual collection-window costs are kept separate. Identifiable source records remain private.

Firecrawl

Firecrawl captured pages from 559 of 607 difficult-site domains

A verified qualification round retained pages from 92.1% of its Firecrawl queue and added verified website evidence for 395 businesses. That is a concrete contribution to the research: businesses that would otherwise have had less website evidence could proceed with a better-supported profile.

This edition describes the 14–19 August qualification round, whose parent population was 4,703 prospects. In that executed round, ordinary retrieval and candidate verification preceded a difficult-site Firecrawl batch. The wider project contract also specified Firecrawl Batch Scrape for resolved business-owned URLs; the executed round and the broader design should not be confused.

What the difficult-site batch produced

Outcome Count Share of 607 submitted domains
Page captured 559 92.1%
Unresolved retrieval 48 7.9%
Verified website evidence added 395 businesses Kept as a business outcome
Captured pages retained without promotion 164 domains Kept as a domain outcome

All 559 HTML/metadata pairs were retained and hash-reconciled. Capturing a page and accepting it for a particular business were separate steps: the first is a retrieval result; the second makes the material useful to the business profile.

What was saved, and what was extracted from it

Evidence or field group Role in the process
Raw HTML and paired metadata Saved page evidence for repeatable extraction
Requested and final URL Records the destination and redirects
Retrieval state, time and provider Identifies the collection outcome
Page hash Detects identical content and links the assessment to its source
Visible service wording and page content Explains the business and supports content-depth assessment
Structured data and page metadata Supplies machine-readable business information
Website links Supplies social-account and relevant internal-page candidates
Public email and telephone links Supplies contact evidence where present
Forms and technical page features Supports the website assessment after extraction

The first four groups describe retrieval evidence. The later groups describe material extracted from the captured pages; they are not all direct Firecrawl response fields. The project’s retrieval contract requested contact-bearing content, including headers and footers, so useful business information was not discarded with a main-text-only view.

It recovered more than obvious access blocks

Firecrawl recovered 24 of 37 cases previously labelled as TLS failures and 11 of 14 labelled as other transport failures. Those results show why a failed ordinary request was not sufficient reason to abandon the business’s website evidence.

The round also exposed a cost problem: a gateway timeout could leave a provider job running. Retrying without reconciling that job risked another charge. Payloads reported 559 credits; the balance showed 619 spent at the last observation, with an estimated 660–700 after outstanding jobs. Small batches of 10–15 URLs were the reliable operating size observed under that round’s 90-second gateway timeout.

The practical lesson is to save the provider job reference before polling, reconcile a timeout against that job, then decide whether a retry is needed. That protects the spend and preserves pages already collected.

How retained pages paid for more than one analysis

The round stored 3,510 HTML occurrences across direct retrieval, candidate verification and Firecrawl, representing 3,478 distinct pages after content deduplication. Those saved pages supported website features, service wording, contact extraction and social-link discovery. The same source could be analysed again without paying to collect it again.

The wider programme recorded 6,720 Firecrawl credits, with a reported cost of roughly £30/$40. Those programme figures are separate from this named batch, and credits are not requests.

Sources and reading the figures

These figures come from retained project receipts and outputs. Field counts describe populated values in returned records, including repeated observations; they are not counts of unique businesses. A returned error item is reported separately from a returned profile. Blank fields are not converted into zero. Completeness flags, ownership decisions and qualification scores are derived after collection, not supplied as scores by the scraper. No additional collection was performed for this edition.

Source material: September provider-consumption receipts, platform raw-record exports in the combined research database, and the verified August qualification closeout where explicitly stated. Programme-wide reported costs and individual collection-window costs are kept separate. Identifiable source records remain private.

Exa

Exa searched 687 missing-site records in under four minutes

In the verified August qualification round, Exa processed all 687 remaining discovery records in 236.2 seconds at a recorded cost of $4.809. It returned a high-confidence or plausible website candidate for 598—87.0%.

The job was specific: find candidate websites for prospects whose initial record lacked a captured website. Google outbound-link resolution ran first. Parallel’s free route completed 40 searches before its fair-usage throttle interrupted that approach; Exa handled the remaining 687 records.

What the search produced

Candidate outcome Records Share of 687
High-confidence candidate 479 69.7%
Plausible candidate 119 17.3%
Directory or other result 75 10.9%
No owned candidate 14 2.0%

The useful output was a candidate destination linked to the business being researched, followed by a decision about relevance. A directory result could help identify the business without becoming its official website. A likely domain then went through page capture and identity verification.

Search and qualification did different work

Step Output used by the next step
Missing-site input Business identity and available location evidence
Exa search Candidate website destinations
Candidate assessment High-confidence, plausible, directory/other or no-owned-candidate classification
Page capture Retained content from the candidate destination
Identity verification Accepted owned-site evidence linked to the prospect
Qualification Website assessment and useful extracted business information

Across the missing-site discovery and verification stage, 258 of the original 1,320 prospects without a captured website gained verified owned-site evidence. That is a result of the joined stage, not 258 sites that can all be attributed to Exa alone.

The distinction is useful to anyone building prospecting automation. Search expands the options quickly; retained page evidence turns an option into something a business report can use. Making these outputs explicit allows search performance and website qualification to be assessed separately.

The named Exa batch averaged about $0.007 per processed record. The wider programme’s reported Exa cost was $22.12. These figures describe different scopes; neither is a current price quote.

Sources and reading the figures

These figures come from retained project receipts and outputs. Field counts describe populated values in returned records, including repeated observations; they are not counts of unique businesses. A returned error item is reported separately from a returned profile. Blank fields are not converted into zero. Completeness flags, ownership decisions and qualification scores are derived after collection, not supplied as scores by the scraper. No additional collection was performed for this edition.

Source material: September provider-consumption receipts, platform raw-record exports in the combined research database, and the verified August qualification closeout where explicitly stated. Programme-wide reported costs and individual collection-window costs are kept separate. Identifiable source records remain private.

Website evidence and direct retrieval

6,131 distinct website captures

The combined evidence register records 6,131 distinct HTML content hashes matched to 6,562 business records. Shared captures can support more than one business record. This is the combined collection across capture routes, not a direct-retrieval-only total or a fresh physical-file count.

August direct-retrieval campaign: 2,383 domains

The documented August campaign attempted 2,893 canonical domains and retained HTML from 2,383—82.4%. It spent zero provider credits. The saved content supported website assessment, contact extraction and discovery of social-account links.

Campaign result Recorded quantity
Canonical domains attempted 2,893
Domains with HTML retained 2,383
Domains in the remaining retrieval queue 510
Businesses scored in the campaign 2,425
Raw social links found 5,423
Canonical account candidates after link reconciliation 4,008
Businesses with those social candidates 1,625

The number of businesses scored differs from the number of domains: business records and domains are different units. Similarly, 5,423 links became 4,008 candidate accounts because repeated links and URL variants needed reconciliation.

One saved page supplied several kinds of evidence

The extraction examined visible text, links, structured JSON-LD, metadata, contact links and forms. It supplied service wording, website contact routes, social-account candidates and technical features used by the assessment. Retaining the original page meant those outputs could be checked and recalculated without another collection request.

For an automation builder, the important design choice is to save the source once and extract several useful outputs from it. The page becomes more valuable than a single score: it can explain what the business does, where the customer can go next and which social accounts should be collected in the next phase.

The campaign’s remaining retrieval queue showed the limit of ordinary fetching. The Firecrawl edition explains the later difficult-site work and its verified results. Zero provider credits describes this retrieval lane’s metering, not zero labour or infrastructure cost.

Sources and reading the figures

These figures come from retained project receipts and outputs. Field counts describe populated values in returned records, including repeated observations; they are not counts of unique businesses. A returned error item is reported separately from a returned profile. Blank fields are not converted into zero. Completeness flags, ownership decisions and qualification scores are derived after collection, not supplied as scores by the scraper. No additional collection was performed for this edition.

Source material: September provider-consumption receipts, platform raw-record exports in the combined research database, and the verified August qualification closeout where explicitly stated. Programme-wide reported costs and individual collection-window costs are kept separate. Identifiable source records remain private.

X

X supplied account history as well as the latest post

The retained receipts record 3,049 X requests against 1,622 submitted candidates. The export contains 1,478 result records, with a resolved user identity in 1,386 and a latest-post date in 1,298.

X used its API route rather than the Apify tasks. The output combined profile information and public metrics with the latest recorded post. It also kept reply, repost and quote flags, which help distinguish original business material from other activity.

What the records contain

Captured field Records with a value / returned records What it contributes
Resolved user identity 1,386/1,478 Account identity
Display name 1,386/1,478 Business matching
Description 1,277/1,478 Services and positioning
Website/account link 1,273/1,478 Profile destination
Account creation date 1,386/1,478 Account age
Followers 1,386/1,478 Audience
Following 1,386/1,478 Profile context
Total posts 1,386/1,478 Accumulated activity
Listed count 1,386/1,478 Platform context
Protected flag 1,386/1,478 Access state
Latest post date 1,298/1,478 Recency
Latest post text 1,298/1,478 Message
Likes 1,298/1,478 Response
Replies 1,298/1,478 Discussion
Reposts 1,298/1,478 Redistribution
Quotes 1,298/1,478 Quoted sharing
Impressions 1,298/1,478 Returned exposure value
Reply/repost/quote flags 1,298/1,478 Content attribution

Of the 1,386 records with a resolved account, 1,298 also contain a latest-post date: 93.7%. This makes recency a useful companion to the audience figure in this returned record set. Description text is available in 1,277 records, providing wording to compare with the website and sector label.

In the electrician sample, five of eight accepted X accounts had latest recorded posts more than a year old. That points to an existing-profile maintenance issue, rather than evidence that electricians should rush to add another X account. An owner can decide whether to refresh the profile or concentrate current project updates on a channel they already maintain.

The reported programme cost was $19.38. Requests, submitted candidates and accepted business accounts are different units; they are shown separately. Counts for likes, replies and impressions describe the returned latest post, not a campaign average or a rate of winning enquiries.

Aggregate field catalogue · Electrician platform comparison

Sources and reading the figures

These figures come from retained project receipts and outputs. Field counts describe populated values in returned records, including repeated observations; they are not counts of unique businesses. A returned error item is reported separately from a returned profile. Blank fields are not converted into zero. Completeness flags, ownership decisions and qualification scores are derived after collection, not supplied as scores by the scraper. No additional collection was performed for this edition.

Source material: September provider-consumption receipts, platform raw-record exports in the combined research database, and the verified August qualification closeout where explicitly stated. Programme-wide reported costs and individual collection-window costs are kept separate. Identifiable source records remain private.

YouTube

YouTube revealed the difference between having a channel and having a video library

The retained receipts record 820 submissions and 1,956 YouTube quota units. The export contains 766 result records: 616 with a resolved channel and 582 with a latest-video date.

YouTube used the Data API route. Its contribution was different from a social-profile check: it supplied the size of the video library, accumulated channel views and the latest video’s title, date and response. These fields let an analyst distinguish a channel page from visible work a customer can actually watch.

What was captured

Captured field Records with a value / returned records What it contributes
Resolved channel identity 616/766 Channel matching
Channel title 616/766 Business presentation
Channel description 478/766 Services and content focus
Subscribers 616/766 Audience
Subscriber-hidden flag 616/766 Interpreting audience availability
Total videos 616/766 Content library
Total views 616/766 Accumulated channel viewing
Channel creation date 616/766 Channel age
Latest video date 582/766 Recency
Latest video title 582/766 Subject matter
Latest video views 582/766 Viewing of that video
Latest video likes 569/766 Response
Latest video comments 546/766 Discussion
Short-format flag 582/766 Format context
Country 316/766 Geographic context

Of the 616 resolved channel records, 582 contain a latest-video date—94.5%. Latest-video views are available for all 582, while likes are available for 569 and comments for 546. Showing those denominators avoids treating the absence of a response field as zero engagement.

A channel’s total views and its latest video’s views answer different questions. The first describes accumulated viewing across its history; the second describes one piece of material. The report keeps them separate so a large historic library cannot silently become a claim that the business is publishing actively now.

For an electrical contractor, video can demonstrate a completed job or explain a service. The dataset supports comparing how businesses present that material through titles, descriptions and dates. The electrician sample itself contains only three accepted YouTube accounts, so it supports a description of those accounts rather than a sweeping recommendation for the whole trade.

The programme recorded no YouTube cash charge. Quota units still measure API consumption; they are not a dollar amount or a count of videos. Subscriber-hidden state is retained alongside subscriber values and should be respected before making audience comparisons.

Aggregate field catalogue · Electrician platform comparison

Sources and reading the figures

These figures come from retained project receipts and outputs. Field counts describe populated values in returned records, including repeated observations; they are not counts of unique businesses. A returned error item is reported separately from a returned profile. Blank fields are not converted into zero. Completeness flags, ownership decisions and qualification scores are derived after collection, not supplied as scores by the scraper. No additional collection was performed for this edition.

Source material: September provider-consumption receipts, platform raw-record exports in the combined research database, and the verified August qualification closeout where explicitly stated. Programme-wide reported costs and individual collection-window costs are kept separate. Identifiable source records remain private.

Companies House

Companies House linked 261 prospects to exact company records

All 4,703 prospects in the August pilot were processed; 261 matched the approved normalised-name-and-postcode rule. Every matched company also had people-with-significant-control evidence in the retained source.

Pilot outcome Records
Prospects processed 4,703
Exact automatic company matches 261
Unmatched under that rule 4,442
Matched companies with PSC evidence 261

The automatic match rate was 5.55%. Its purpose was to establish an exact company link, not to judge the proportion of local businesses that were incorporated. The basic-company snapshot was dated 1 August and the PSC snapshot 14 August 2026.

What the match added

The pilot connected a prospect identity and postcode to an authoritative company record, then retained matching status and PSC coverage. Its output was a separate company-enrichment workbook with matched, unmatched and review material. It was not silently joined into every business row in the qualification master.

For a practitioner, this is a distinct stage from searching for a trading website. A trading name can lead to the correct website while still requiring a separate legal-identity match. The retained exact matches provide that corporate context without changing the website and social observations.

For the case studies, company-age or size comparisons must use an established company link or a clearly labelled platform-reported field. Account creation date is a different measure. Keeping those distinctions in the source design prevents a plausible-sounding story about younger firms from being built on the age of a social profile instead.

Source: the August qualification closeout and the Companies House pilot receipt, including reconciliation of all 4,703 dispositions and the 261 exact matches. Identifiable company and PSC records are not included in this reader edition.

Sources and reading the figures

These figures come from retained project receipts and outputs. Field counts describe populated values in returned records, including repeated observations; they are not counts of unique businesses. A returned error item is reported separately from a returned profile. Blank fields are not converted into zero. Completeness flags, ownership decisions and qualification scores are derived after collection, not supplied as scores by the scraper. No additional collection was performed for this edition.

Source material: September provider-consumption receipts, platform raw-record exports in the combined research database, and the verified August qualification closeout where explicitly stated. Programme-wide reported costs and individual collection-window costs are kept separate. Identifiable source records remain private.

Sources are the retained project editions and their receipts. Named batches, programme totals and record types retain their original scope. Nothing was collected again for this prototype.