---
name: linkedin-stats
description: Scrape your own LinkedIn posts and engagement (likes, comments, reposts, impressions, who reacted and commented) into local JSON using Claude in Chrome, then build a private dashboard. Use for "analyse my LinkedIn", "refresh my LinkedIn stats" or a monthly update.
---

# LinkedIn scrape playbook

> **Validated 22 Sep 2026:** the posts endpoint below still returns posts and counts. Keep the LinkedIn tab as the visible, front tab: since LinkedIn's September 2026 front-end change, a background LinkedIn tab freezes and every call times out.

Scrapes the logged-in user's own posts and engagement via LinkedIn's internal
voyager API, replayed **in-page** with `javascript_tool` so credentials never
leave the browser. Run from this skill's folder. Produces `data/posts.json`, `data/reactors.json`,
`data/commenters.json`. Then run `node src/analyze.mjs && node src/build-dashboard.mjs`.

## Method (validated 2026-07-18)

1. **Setup**: load claude-in-chrome tools, `tabs_context_mcp`, navigate to
   `https://www.linkedin.com/in/me/recent-activity/all/`, wait 4s.

2. **Discover the posts endpoint** (queryId hash changes over time — never hardcode):
   scroll the page 2-3×, then in-page:
   `performance.getEntriesByType('resource').map(e=>e.name).filter(u=>u.includes('voyagerFeedDashProfileUpdates')).pop()`
   Keep that full URL as the template. NOTE: the extension blocks returning
   query strings/cookies from javascript_tool — parse/use URLs **inside** the
   page, return only derived scalars.

3. **Paginate**: plain `start` is IGNORED — the server returns page 1 forever.
   Thread `metadata.paginationToken` from each response into the next request:
   insert `paginationToken:<encodeURIComponent(token)>,` right after `variables=(`.
   Headers for all voyager fetches:
   `csrf-token` = JSESSIONID cookie value (strip quotes),
   `accept: application/vnd.linkedin.normalized+json+2.1`,
   `x-restli-protocol-version: 2.0.0`.
   Loop until no/unchanged token (~116 pages for full history). Run as a
   background async loop writing to `window.__lsm` with a status object —
   javascript_tool times out at 45s, so never await long loops in the tool call.

4. **Extract per page** from `included`:
   - `feed.Update`: activity id from entityUrn, `commentary.text.text`,
     content keys, `actor.name.text`
   - `social.SocialDetail`: entityUrn activity id → `*totalSocialActivityCounts`
     ref + `threadUrn` (needed for reactions; sometimes a ugcPost urn)
   - `feed.SocialActivityCounts`: numLikes/numComments/numShares,
     `numImpressions` (only recent ~440 posts), reactionTypeCounts
   - **Timestamp**: `Number(BigInt(activityId) >> 22n)` = epoch ms (works back to 2013)
   - **Reshare detection**: `isReshare = actor.name.text !== '<your display name>'` (read your name from the profile card on the page first).
     `resharedUpdate` is always absent in this format — do NOT trust it. Reposted
     items carry the ORIGINAL author's social counts (viral-outlier trap).
   - **CRITICAL — counts join (the ugcPost trap):** a post's SocialActivityCounts
     is sometimes keyed on `urn:li:ugcPost:<n>` or `urn:li:share:<n>`, NOT
     `urn:li:activity:<n>`. Videos ALWAYS are; a large fraction of images/text are
     too. If you only match the activity-keyed SAC you'll silently record 0
     likes / 0 comments / null impressions for ~30-40% of posts (a real 867-like
     post came back as 0). Do NOT trust `numLikes` from `profileUpdates` alone.
     The reliable source of post-level counts is
     `/voyager/api/feed/updates/urn:li:activity:<id>` → in `included`, take the
     `SocialActivityCounts` whose entityUrn matches
     `socialActivityCounts:urn:li:(ugcPost|activity|share)` and is NOT a
     `:comment:` one (pick the max-count one if several). It carries
     numLikes/numComments/numShares/numImpressions authoritatively. Either fetch
     counts this way for every post, or scrape fast via profileUpdates then
     re-fetch counts per-post to correct. Sanity check: any 2024+ post with
     0 likes AND 0 comments AND null impressions is almost certainly a join miss.
   - Note: for videos, `numImpressions` exists on the ugcPost SAC (LinkedIn also
     tracks video *views* separately, which this field is not).

5. **Reactions** (who liked): classic REST endpoint, no queryId needed:
   `/voyager/api/feed/reactions?count=50&q=reactionType&start=N&threadUrn=<enc(threadUrn)>`
   `included` has `feed.social.Reaction` (reactionType, actorUrn) +
   `identity.shared.MiniProfile` (firstName, lastName, occupation,
   publicIdentifier); join `Reaction.actorUrn === MiniProfile.entityUrn`.
   Cap at 20 pages (1,000 reactors) per post — LinkedIn returns 1st-degree
   connections first, counts are already exact from SocialActivityCounts.

6. **Comments**: `/voyager/api/feed/comments?count=50&q=comments&sortOrder=CHRON&start=N&updateId=activity:<id>`
   `feed.Comment` (commenterProfileId, commentV2.text, createdTime) joined to
   MiniProfile by trailing id of entityUrn.

7. **Engagement harvest**: only for `!isReshare` posts with likes/comments > 0,
   newest first. 300-500ms jittered sleeps, backoff 8s on 429, retry ×3.
   ~2,600 requests ≈ 2h single loop; run 2 concurrent workers over
   disjoint slices (newest half + oldest half) ≈ 1 req/s combined. Don't go
   wider — rate-limit risk. Give each worker its own status object and a stop flag.

8. **Getting data out** (avoids blowing context): click a neutral page spot to
   focus the document (clipboard needs focus), then
   `navigator.clipboard.writeText(JSON.stringify(...))` in-page, then Bash
   `pbpaste > data/xxx.json` (macOS; on Linux use `xclip -o` or `wl-paste`) and validate with node. 1.6MB works fine.
   Dedupe reactors on (postId,key) and comments on (postId,key,ts) at save time.

9. Update `data/meta.json` (scrapedAt, followers from the profile card) and
   rebuild: `node src/analyze.mjs && node src/build-dashboard.mjs`.

## Gotchas
- javascript_tool CDP timeout is 45s; background loops keep running after it — poll status.
- The MEMBER_SHARES feed contains ~25% reposts; always separate them.
- `performance` resource buffer overflows silently; grab the template URL early.
- Reactor identity fetch beyond ~page 20 adds little (long-tail 2nd/3rd degree).
