Infographic summarising 5.2 million API requests from 389 accounts: 82% used multiple platforms, with AI supporting creator analysis, discovery, social listening and brand voice.

What 5.2 million data requests reveal about how 389 businesses are using social data with AI

Most people still think about social media data in terms of posts, likes and follower counts.

But a public post is only the smallest unit in a much richer dataset. Around it sits a profile, a publishing history, a network of replies, a changing set of engagement signals, and often a video, transcript or comment thread that gives the post its real meaning. Follow those signals over time and social media starts to look less like a content feed and more like a continuously updating record of what people care about.

AI makes that record far more usable.

Models such as Claude can already read hundreds of posts, group recurring complaints, compare a creator’s recent work with their usual performance, or identify the language customers use when they are considering a competitor. The interesting shift is not that AI can generate another social post. It is that AI can help us interpret the enormous amount of public evidence people are already creating.

I see this firsthand through SocialCrawl, a social-media data API. I recently looked at a 91-day anonymised production window containing 5,217,266 requests from 389 accounts. This is behavioural data rather than a survey, so it cannot tell us everything about the intent behind each request. An account may represent one person, a team or a product used by many people.

Even with those caveats, the pattern was striking: 319 of 389 accounts, or 82%, used data from at least two platforms. The average paid account touched 5.29 platforms and 19.1 different endpoints.

People were not simply checking how one account was performing. They were combining many kinds of social data to answer much more interesting questions.

Social data is much richer than the feed suggests

Consider what can be publicly observable around a single piece of content:

  • The post itself: text, caption, images, video or audio.
  • Its context: author, publication time, hashtags, mentions and links.
  • Its reception: views, likes, comments, shares and replies, where the platform makes them available.
  • The conversation underneath it: questions, disagreements, jokes, requests and follow-up experiences.
  • The author’s wider history: topics, formats, posting cadence, partnerships and changes in audience size.
  • Its evolution: what happened in the first few hours, what continued to travel and what disappeared.
  • Its relationship to other posts: repeated phrases, copied formats, common complaints and ideas crossing between communities.

No single field is definitive. A view on one platform may not mean the same thing as a view on another. Some metrics are unavailable for particular formats. A loud conversation may be influential without being representative.

But taken together, these signals can describe more than popularity. They can reveal emerging language, shifting preferences, unmet needs, creative conventions and the difference between what people say they like and what they actually respond to.

The biggest gains often come from enrichment. A search result tells you that a conversation exists. The post, transcript and replies tell you what the conversation is actually about. A follower count gives you scale. Publishing history and commercial results help you judge whether that scale matters.

AI is useful because it can work across those layers. It can turn a pile of records into a set of possible patterns—provided the original evidence remains attached.

One question rarely lives on one platform

The fact that 82% of the accounts in the sample used more than one platform matters.

Communities do not distribute themselves neatly. A product problem might first surface in a Reddit thread, get reframed in a TikTok video, attract practical questions on YouTube and become a more polished industry opinion on LinkedIn. Each surface provides a different lens on the same underlying subject.

That does not mean every project needs every network. It means the question should determine the sources.

If you are researching technical objections, forums and long-form discussions may be more valuable than high-reach short-form content. If you are looking for a new creative format, TikTok, Instagram Reels and YouTube Shorts may matter more. If you are studying how a company speaks in public, posts and replies from the company’s own accounts are the relevant dataset.

The real opportunity is not “monitor everything.” It is to combine the right public evidence for the decision you are trying to make.

Four patterns in how people are using social data

1. Creator analysis is moving beyond follower counts

One recurring pattern is the continuous tracking of a known creator roster. One anonymised customer follows roughly 200 creators, regularly collecting their latest non-pinned Reels.

The useful output is not a leaderboard of who has the most followers. It is a history of what each creator publishes and how that work performs over time.

That history lets a team ask better questions:

  • Is this creator improving, declining or simply experiencing normal variation?
  • Does one viral post distort their average?
  • Which content performs unusually well for this creator’s typical audience?
  • Do comments suggest genuine interest, confusion or a conversation unrelated to the campaign?
  • Does reach translate into clicks, revenue or another result the business actually values?

This is where first-party data and public social data become especially powerful together. Public data can show the content, audience and visible response. A company’s own affiliate, conversion or revenue data can show the commercial outcome.

AI can then help explain anomalies: why a post may have travelled, which themes recur in the strongest work, or how the response differs from the creator’s baseline. It should not make a renewal decision from a single post. The job is to give a person a clearer comparison and the evidence behind it.

2. Creator and content discovery is becoming continuous research

Another pattern begins without a fixed roster. The question is closer to: Who is already making credible content about this subject? Which voices are emerging in this category? What format is beginning to spread?

One customer in the sample scanned 23,116 public handles in a day as part of a location-filtering workflow. The number is less interesting as a feat of scale than as evidence of a different approach to discovery.

Discovery is no longer limited to browsing a feed, searching a hashtag and saving a few names. A broad set of public candidates can be collected, deduplicated and narrowed using observable criteria such as location stated in a profile, language, subject relevance, posting recency or content format.

AI is most helpful after that first narrowing. Give a model a manageable group of candidates and it can explain why each person appears to fit a brief, identify possible conflicts and point to the posts that support its view.

The evidence requirement is important. A model should not infer a person’s location, audience or commercial performance from their appearance or username. “There is not enough evidence” is a valid and useful result.

The same workflow extends beyond influencer marketing. It can be used to find subject-matter experts, customer advocates, niche communities, potential podcast guests, emerging creative formats or people documenting an experience relevant to a research question.

3. Social listening is moving beyond a sentiment score

Traditional social listening often compresses a large conversation into a chart of positive, negative and neutral mentions. That may be tidy, but it removes much of the information that makes the conversation valuable.

A more useful project starts with a specific question:

What are freelance designers saying about the onboarding experience for product X, and what evidence do they give for their opinion?

That question contains an audience, a subject and a request for evidence. It creates a much better AI task than “summarise what people think of product X.”

The workflow is also revealing. In the 91-day sample, broad cross-platform search represented only 0.58% of requests but 11.5% of credits. Comment and reply extraction accounted for 23% of all credits. Credit consumption is not the same thing as user value, but it does suggest that meaningful listening often involves more than finding the first mention. People are spending substantial effort retrieving the conversation around the content.

Once that evidence is available, AI can group posts into themes, distinguish questions from complaints, surface contradictions and identify the language that appears repeatedly. Every finding should link back to the source posts that produced it.

This creates a wide range of useful workflows:

  • Product research: find recurring friction, workarounds and feature requests in people’s own words.
  • Launch analysis: separate confusion, curiosity, praise and genuine failure reports during the first days of a launch.
  • Sales research: build an evidence-backed library of objections, comparisons and “switching from” conversations.
  • Community research: compare how experts, practitioners and broad consumer audiences discuss the same subject.
  • Customer support: identify clusters of public reports that may point to a shared issue.
  • Creative research: find recurring hooks, phrases, formats and audience questions without separating them from their original context.
  • Market intelligence: observe where competitor comparisons are appearing and which trade-offs people mention without being prompted.

The output is not a universal measure of public opinion. Social data has selection bias, platform bias and plenty of noise. Its strength is different: it offers timely, specific and often highly contextual evidence about what people are choosing to discuss in public.

4. Brand voice can be studied as behaviour, not a vibe

Brand-voice analysis is a smaller pattern in the production data, but it is an unusually good fit for AI.

Most brand guidance lives in a static document: a handful of adjectives, some approved examples and a list of things not to say. The actual brand voice lives in hundreds of posts and replies written by different people under different pressures.

Public social data makes that behaviour observable. A team can study repeated openings, overused calls to action, changes in tone, reply speed, differences between platforms and the gap between published guidance and real output.

AI can compare a new post with a set of approved examples and flag a possible mismatch. But a useful flag needs to be concrete. It should identify the exact phrase that triggered the concern, show the examples used for comparison and suggest an alternative for an editor to accept, change or reject.

That feedback loop is more valuable than a mysterious “brand score” of 72%. Over time, the editor’s decisions also create a better record of what the organisation actually considers on-brand.

The emerging workflow: retrieval, enrichment, analysis, judgment

Across these use cases, the same basic shape appears:

Start with a real question → collect relevant public records → add the context that matters → compare like with like → ask AI to find and explain patterns → inspect the sources → make a decision.

The model sits between the data and the decision. It is an analyst, not an oracle.

This distinction changes the quality of the output. If a model is asked to “monitor the internet” without a defined source set, it can give a fluent answer without giving a trustworthy one. If it receives a bounded collection of posts, comments and time-stamped metrics, it can perform a much more useful task.

A good prompt for Claude might look like this:

Review the attached social posts and comments for recurring themes related to onboarding.

For each theme:
- describe the pattern in plain language;
- cite the IDs of the posts that support it;
- include one short piece of supporting evidence;
- note any examples that contradict the pattern;
- state whether the evidence is strong, mixed or insufficient;
- suggest one question a researcher should investigate next.

Do not make claims about the wider market. Base every observation only on the supplied records.

The prompt is not technically complex. Its strength comes from asking the model to show its work, respect the limits of the dataset and preserve uncertainty.

What becomes possible when social data is treated as a research layer

Once social media is treated as structured evidence rather than an endless feed, a number of new combinations become possible.

A product team can compare support tickets with public workarounds. A researcher can connect interview themes with the language appearing in online communities. A brand can see not only which creative performed, but what viewers asked beneath it. A sales team can compare the objections in call notes with the ones prospects express publicly. A founder can follow a niche topic across communities without relying on whichever post an algorithm happens to put in front of them that morning.

None of this requires a fully autonomous agent. In fact, the most useful early workflows are usually narrow and repeatable. Pick one question that returns every week, gather the same kinds of evidence each time, and let AI help with the reading and comparison.

The technology is making collection and analysis easier. The harder—and more interesting—work is deciding what to observe, which context changes the meaning, and what evidence is strong enough to act on.

The broader lesson from 5.2 million requests

The main takeaway from this production window is not that people want a bigger social dashboard.

It is that social media is becoming part of the research stack.

People are using it to study creators over time, discover relevant voices, understand the conversation beneath a post and make qualitative ideas such as brand voice more observable. They are combining platforms because no single network contains the whole story. And they are going deeper than the first search result because the surrounding comments, transcripts, history and changing metrics are often where the insight lives.

AI does not make the underlying data complete or unbiased. What it changes is the amount of evidence a person can realistically examine. It lets a small team read more broadly, compare more consistently and ask better follow-up questions.

The social web has been producing this dataset in public for years. We are only beginning to learn how to use it.

Share this post!
Selene Lee
Selene Lee
Articles: 37

Leave a Reply

Your email address will not be published. Required fields are marked *