All tags

HOME
AI Company News Op-Eds OSINT OSINT Case Study OSINT Events OSINT News OSINT Tools Press Release Product Updates SL API SL Crimewall SL Professional for i2 SL Professional for Maltego Use Сases

Social Network Analysis: Reading the Structure

In late 2001, Enron's collapse triggered a federal investigation that eventually put the company's internal email archive, more than half a million messages exchanged between employees over several years, into the public record. Read one email at a time, it was an unremarkable stream of corporate correspondence. Mapped as a communication network, a different structure emerged: researchers later found that during the company's crisis period, the network grew denser and more centralized than usual, with senior executives forming a tight, mutually reinforcing cluster that brokered nearly every connection to the rest of the company.

That structure wasn't visible in any single email; it only became visible once the relationships between senders and recipients were mapped, measured, and interpreted as a network. The same principle holds whether the entities involved are employees exchanging email or accounts interacting on social media: raw relationship data says very little on its own, and turning it into something defensible takes a disciplined process, a clearly defined question, carefully built connections, calculated metrics, and findings checked back against the underlying evidence before anyone treats a pattern as a conclusion.

In this article, we examine the workflow that turns raw relationship data into a defensible finding, what a centrality score actually can and can't tell an analyst, the mistakes that undermine otherwise solid analysis, and the data-quality and access limitations every social network analysis has to reckon with.

What Is Social Network Analysis?

Social network analysis studies the relationships between entities, called nodes, rather than treating each entity as an isolated data point. Connections between nodes, called edges, can be directed or undirected and weighted by how strong or frequent the relationship is. A handful of centrality metrics, degree, betweenness, closeness, and eigenvector centrality, help analysts identify which nodes matter structurally, and community detection reveals clusters that interact more with each other than with the network at large.

Beyond the Score: What Centrality Doesn't Tell You

A centrality score is a measurement of a node's position within a specific, already-collected dataset. It is not a measurement of real-world importance, and treating the two as the same thing is one of the most common ways a solid analysis produces a wrong conclusion.

Three things sit between a high centrality score and an actual finding. The score only reflects what was collected: an account with genuinely more real-world influence than anyone in the dataset will show low centrality if it simply wasn't captured. The score also doesn't know what an edge represents: a network built from mentions and a network built from direct messages will rank the same accounts completely differently, because the underlying relationship means something different in each case.

Structural position and actual behavior aren't the same claim, either. In the Enron corpus, ranking near the top across multiple centrality measures didn't by itself distinguish an ordinary senior manager from someone later implicated in the fraud; a high score is a structural fact about a communication pattern, not a legal one. Centrality identifies where to look, not what to conclude.

How Social Network Analysis Works

Social network analysis moves through a structured sequence, and skipping a stage tends to produce a graph that looks convincing without actually holding up.

Defining the question and the network boundary comes first. What relationship is actually being analyzed, and which accounts, time period, or platform fall inside that boundary? Setting this after looking at the results, rather than before collecting data, is one of the most common ways an analysis gets shaped to fit a preferred conclusion.

Collecting relationship data draws from sources like social media interactions, public profiles, communication records, group memberships, shared infrastructure, event participation, or existing structured datasets, depending on what the question actually requires.

Creating nodes and edges means deciding exactly what counts as a relationship. A mention and a reply imply different levels of engagement, and treating them as interchangeable edges can distort the resulting network without anyone noticing.

Enriching and normalizing the network resolves duplicate accounts, aliases, and inconsistent identifiers, and attaches relevant attributes to each node, work that determines whether the network actually reflects reality or just reflects how the raw data happened to be formatted.

Calculating metrics and detecting communities applies centrality measures, clustering, and path analysis to the cleaned network, turning the graph into something that can actually be interpreted rather than just looked at.

Visualizing and investigating means filtering the network down to what's relevant, expanding around genuinely interesting nodes, and comparing subgraphs, rather than presenting the entire raw graph as though every node and edge carries equal weight.

Validating the finding means returning to the underlying source evidence behind any significant relationship before drawing a conclusion from it. A structurally central node is a lead. Verified activity behind that node is a finding.

The two terms overlap enough that they're often used interchangeably, and in practice, much of the underlying work is the same. Link analysis is the term more common in investigative and law enforcement contexts, with a stronger emphasis on individual relationships and pivoting from one entity to the next. Social network analysis leans more heavily on structural, mathematical measures applied across an entire network at once, degree, betweenness, density, community detection, rather than tracing one chain of connections at a time.

In practice, a strong investigation usually uses both without worrying too much about which label applies at any given moment: link analysis to build and verify individual connections, and network-level metrics to see where those connections sit relative to the whole.

Common Mistakes in Social Network Analysis

A handful of recurring mistakes undermine otherwise solid network analysis work, often without anyone noticing until a finding falls apart under scrutiny.

Treating every observed connection as equally meaningful flattens a graph that actually contains very different strengths of relationship into a set of visually identical lines.

Ignoring missing data treats an incomplete network as though it were the complete picture, when private accounts, deleted content, and platform access limits routinely leave real relationships out of the dataset entirely.

Defining network boundaries after seeing the result lets the conclusion shape which accounts or relationships counted as in-scope, rather than the other way around. 

Mixing incompatible relationship types without labeling them merges a mention, a reply, and a shared IP address into identical-looking edges, even though each implies a very different kind of connection.

Confusing correlation or proximity with proof treats two accounts appearing in the same cluster as evidence of coordination, when shared timing or topic alone doesn't establish that the accounts are actually connected.

Ignoring time treats a network as a fixed structure, when real relationships form, strengthen, and dissolve, and a static snapshot can misrepresent activity that only overlapped briefly or has since ended entirely.

Limitations, Privacy, and Data Quality

Every social network analysis is built on a dataset that is, to some degree, incomplete. API and platform access limits shape what can be collected in the first place, and those limits have gotten more restrictive in recent years, not less. X's decision to end free academic API access in 2023 is a clear example: the number of published studies using Twitter data dropped 13 percent the following year, a direct consequence of researchers losing access they'd previously relied on.

Private accounts, deleted content, and platform-specific visibility settings all remove real relationships from a dataset without any indication that they're missing. Sampling bias distorts a network in ways that aren't always obvious: a dataset built from one collection method or one time window can look structurally different from the same underlying network sampled another way. Duplicate and false identities, bots, sockpuppets, and reused aliases, can inflate or distort node counts if they aren't identified during enrichment.

Ethical and privacy obligations don't disappear because data is technically accessible, and network structure changes over time in ways a single snapshot won't capture. All of this is why source-level validation matters more than the graph itself: the network is a starting point for investigation, not the finished product.

Practical Applications Worth a Closer Look

Criminal or fraud networks connect people to accounts, accounts to companies, and companies to the shared identifiers and intermediaries that tie an otherwise fragmented set of entities together, often surfacing intermediaries that wouldn't stand out in any single record, the same kind of structure that made Enron's brokered executive cluster visible only once mapped.

Coordinated influence and disinformation networks look unremarkable account by account, but reveal distinct operational roles once interaction patterns are mapped and checked against actual posting behavior.

Threat-actor ecosystems map aliases to infrastructure and infrastructure to online communities, frequently revealing the bridge entities that connect campaigns which would otherwise look unrelated.

The Takeaway

Social network analysis works because it studies relationships instead of isolated entities, and relationships are where coordination, influence, and hidden structure actually live. But a graph and a finding are not the same thing. A defined question, a carefully built network, and validation against real evidence are what separate a rigorous analysis from an attractive picture that happens to have some numbers attached.

The strongest social network analysis work treats centrality scores and detected communities as leads worth checking, not conclusions that speak for themselves, a discipline that matters as much as any metric or tool.

FAQ

They overlap substantially. Link analysis emphasizes tracing individual relationships between specific entities, while social network analysis applies structural, mathematical measures like centrality and clustering across an entire network at once.

How do you validate a social network analysis finding?

By returning to the underlying source evidence behind any structurally significant node or connection before treating it as a conclusion, since centrality scores and detected communities are leads, not proof, on their own.

Why can a centrality score be misleading?

Because it only reflects the dataset that was actually collected, and it doesn't capture what the underlying relationship means. An incomplete dataset or a poorly defined edge type can produce a high score for a node that isn't actually significant.

What data-quality issues affect social network analysis the most?

Missing data from private accounts and platform access limits, sampling bias from how data was collected, and duplicate or false identities that inflate node counts if they aren't resolved during enrichment.

When should network boundaries be defined in an investigation?

Before data collection begins, based on the specific question being asked. Defining or adjusting boundaries after seeing preliminary results risks shaping the analysis to fit a preferred conclusion rather than the actual data.


Want to see social network analysis applied with full source traceability? Book a personalized demo with one of our specialists and discover how SL Crimewall helps investigators build, validate, and document network findings without losing sight of the evidence behind every connection.

Share this post

You might also like

You’ve successfully subscribed to Social Links — welcome to our OSINT Blog
Welcome back! You’ve successfully signed in.
Great! You’ve successfully signed up.
Success! Your email is updated.
Your link has expired
Success! Check your email for magic link to sign-in.