Maritime Attention Index — Methodology v1.0 · 31 July 2026
What the Index measures
The Maritime Attention Index measures how much attention the world is paying to shipping on any given day.
Most indices of attention count coverage — how many articles were published. This one is built primarily on what people did: what they searched for, what they looked up, what they posted about. Coverage is an input to public attention, not a measure of it. A story can be heavily covered and widely ignored.
50 is the baseline. A reading of 100 means roughly twice the baseline attention; 25 means about half.
The Index does not measure sentiment. It says how much attention shipping is receiving, never whether that attention is favourable.
The four signals
| Signal | Weight | What it captures | Source |
|---|---|---|---|
| Search | 30% | People actively looking up shipping | Google Trends |
| Research | 30% | People reading about how shipping works | Wikipedia pageviews |
| News | 20% | Press output about the industry | Global Database of Events, Language and Tone (GDELT) |
| Social | 20% | Public conversation | X and Bluesky |
Why these weights
The weights are a claim about what the Index measures.
Search and research carry the most because they are the strongest evidence of genuine engagement. Someone typing "container ship" into Google, or opening the Wikipedia article on bulk carriers, has actively gone looking. Nobody does either on a schedule, and neither can be manufactured by a press office.
News is deliberately not the heaviest signal. It counts articles published, which is press output. It belongs in the Index — coverage both reflects and drives attention — but an index whose central claim is "engagement, not coverage" cannot be led by a coverage count.
Social sits alongside news rather than above it, despite being closer to genuine engagement. Organisations post automatically: a news outlet's every article generally carries at least one post. News and social therefore partly share a cause, and weighting social heavily would count the same underlying event twice.
How each signal is built
Every signal follows the same principle: measure the maritime series against a comparable control, not against a raw total. Each platform has its own long-run drift — traffic falling, source lists changing, user bases growing — and dividing by a raw total mistakes those platform changes for changes in maritime attention. Where a good control exists, it is used.
Search — Google Trends, 30%
Four vessel-type terms, queried as a single union: bulk carrier, cargo ship, container ship, oil tanker.
Why these terms. They are nouns for things that exist whether or not anything is happening to them. Incident words — "Red Sea", "piracy", "collision" — were deliberately excluded: they would manufacture spikes whenever a particular category of incident occurred, turning the Index into a measure of which kind of accident happened rather than how much attention shipping received. Terms like "shipping" and "freight" were also excluded, being heavily contaminated by parcel delivery and unrelated commerce.
Why one query rather than four. Google normalises each request independently, so four separate series have no common scale. Sending them as a union means one normalisation and one interpretable series. This was validated before adoption: across 126 months the union equals the sum of its parts to within 0.6%.
Resolution. Google returns daily data only for short windows and monthly data for long ones. The published daily series is reconstructed from overlapping windows benchmarked to the monthly series, using a method chosen specifically for its behaviour at the edges of events, where a spike is only partly visible to the monthly benchmark.
Research — Wikipedia, 30%
Pageviews for eleven core maritime articles, divided by a control basket of twenty comparable articles — not by total Wikipedia traffic.
Why not total traffic. Wikipedia's overall traffic is dominated by the main page, current events and entertainment, which behave nothing like reference articles. Measured as a share of all Wikipedia traffic, maritime attention appears to have fallen about 40% since 2016. Measured against comparable reference articles, it has risen slightly.
How the control was chosen. 1,603 candidate articles were screened on four gates: enough traffic to be stable, no responsiveness to maritime events, no page moves or title changes, and no unexplained steps of their own. 159 survived. From those, the twenty used are all event-exposed technical and infrastructure topics — power outages, dam failures, level crossings, mid-air collisions — chosen to match the demand type of the maritime articles rather than merely their subject.
That last point matters more than it appears. Purely definitional articles ("what is concrete") have been hit hardest by AI answer engines absorbing lookup traffic. Controlling with definitional articles alone would have flattered the maritime series considerably. Matching on demand type removes about half the apparent outperformance, and is the more conservative reading.
The basket is combined by median, so a single article's page move cannot move the denominator.
News — GDELT, 20%
A count of maritime articles published worldwide each day, divided by English-language article volume.
What is counted. Matching runs against the article headline only. GDELT's metadata field also contains outbound link URLs, and matching the whole field counts a maritime keyword appearing in an unrelated article's link as a maritime article.
Vocabulary. Every keyword must be self-qualifying: a term that is unlikely to mean anything else. Ship types, industry actors, infrastructure and commercial language all count; bare words like "ship" and "vessel" were tested and rejected at roughly 50% precision, pulling in spacecraft, sports idiom, marine archaeology and warehouse software. Naval, cruise and fishing coverage are excluded as different industries.
Coverage. The News signal begins on 1 January 2020, because the headline metadata it depends on is not available earlier. A 33-day window in October–November 2020, when GDELT's crawler dropped roughly a quarter of its sources, is published as no data rather than as a low reading — that is a change in the instrument, not in the world, and no normalisation repairs it.
Social — X and Bluesky, 20%
Posts matching a maritime term basket, divided by posts matching a fixed control basket, on both platforms, combined 90/10.
Why a control basket here too. Both platforms' overall volumes change for reasons that have nothing to do with shipping. The ratio cancels most of it.
Why 90/10. X carries far more maritime conversation. Bluesky is included as a deliberate hedge against single-platform risk rather than as an equal partner. Because Bluesky did not exist for most of the measured period, it is chain-linked: its level is set from its overlap with X rather than from the baseline window directly. Its readings therefore inherit any error in X's baseline.
Retweets are excluded on X, because they are three quarters of all matches and would make the signal a measure of amplification rather than posting. Bluesky does not index reposts, so including them on one platform and not the other would make the two incomparable.
The baseline
50 = the average level of 1 May 2022 – 30 April 2023, for every signal.
This window was chosen because it is the calmest twelve months available: after the container-shipping crisis, before subsequent disruption, and free of the measurement changes that affect other periods. Two of the four signals selected it independently, on the same criterion, without reference to each other.
What "baseline" means here: it is the quietest recent period we can measure, not a long-run historical average. Most other periods therefore read above 50. That is a deliberate choice — anchoring to a busier window would mean burying the very events the Index exists to detect inside its own definition of normal.
A single fixed baseline across all four signals is not cosmetic. The composite is a weighted average of four indices; if each were anchored to a different period, a signal anchored in its own quiet stretch would contribute a systematically higher number than one anchored in a busy stretch, and the declared weights would silently mean something else.
Combining the signals
The composite is a weighted arithmetic mean of the four component indices, each first rescaled onto a common range.
Why the components are rescaled
The four instruments measure genuinely different behaviours, and they arrive on genuinely different natural scales. Over the published series, maritime Wikipedia lookups move across a 1.8× band between their typical low and typical high, while X and Bluesky move across 7.1×:
| Component | Typical range (5th–95th percentile) |
|---|---|
| Social | 7.1× |
| News | 4.7× |
| Search | 3.7× |
| Research | 1.8× |
This is the ordinary problem of combining measurements taken in different units. Sea level rises in millimetres while snowfall accumulates in feet; both are real measures of a changing climate, and comparing them means putting them on a common scale rather than declaring the millimetres unimportant.
Left unscaled, the mismatch does something specific and undesirable. A weighted average gives each component influence proportional to its weight multiplied by how much it moves — so the declared weights become a statement about the level each component contributes rather than its share of the movement. On the raw components that gap was severe:
| Component | Declared weight | Share of movement, unscaled |
|---|---|---|
| Search | 30% | 23% |
| Research | 30% | 8% |
| News | 20% | 18% |
| Social | 20% | 50% |
Research, one of the two heaviest declared weights, drove a twelfth of the Index. Social, one of the two lightest, drove half of it. The published weights did not mean what a reader would reasonably take them to mean.
The rescale
Each component is rescaled around its own baseline level:
y = B × (x ÷ B)^k
where B is the component's average over the baseline window and k is a fixed exponent. Because the calculation divides by B first, a component sitting at its baseline is unchanged — the baseline stays at 50, and only movement away from it is stretched or compressed. The transform is monotonic, meaning no component's shape changes at all: every peak, trough and turning point stays exactly where it was, and only the distance travelled differs.
| Component | B | k | Effect |
|---|---|---|---|
| News | 47.81 | 1.09 | stretched slightly |
| Research | 52.72 | 1.96 | stretched roughly 2× |
| Search | 49.89 | 1.20 | stretched slightly |
| Social | 54.83 | 0.74 | compressed |
The exponents are solved so that all four components end with the same standard deviation — 62.97, the mean of the four unscaled figures, so the components meet in the middle rather than one being made the yardstick. Since influence is weight × movement, equal movement makes influence collapse to weight alone:
| Component | Declared weight | Share of movement |
|---|---|---|
| Search | 30% | 30.0% |
| Research | 30% | 30.0% |
| News | 20% | 19.9% |
| Social | 20% | 20.1% |
What this costs, stated plainly: research's exponent of 1.96 means a genuine 2× move in maritime Wikipedia lookups is published as roughly a 4× move. That is a deliberate change of units, not a measurement. The underlying unscaled values are published alongside the rescaled ones in the data file, so any reading can be traced back.
The exponents and baselines above are frozen. They were derived once, from the 2,343 complete days between 1 January 2020 and 20 July 2026, and are fixed constants in the build. Re-deriving them as new data arrived would let every new day silently rewrite the whole published history.
When numbers are final
The Index publishes every day as soon as it can.
Three of the four signals are final when first published. Wikipedia pageviews, GDELT article counts and the X/Bluesky post counts are all counts of things that already happened, and none of them is revised afterwards.
Google Trends is different. It reports a sample rather than a census, and it resamples: a figure asked for today and the same figure asked for next week can differ slightly. Google settles a calendar month at the end of the following month — so a value published on 3 August is subject to change until 30 September, and fixed from then on.
The Index handles that by publishing and labelling rather than waiting:
- Provisional days are real measurements, published the day they are
available. They are flagged in the published data — provisional in /data/latest.json — rather than annotated on the front page, because the amount they can still move is far too small to qualify the reading a reader is looking at. The figure is in the next paragraph.
- Final days are locked permanently. Once a month freezes, the Index stores
the published figures and republishes them verbatim for ever, even if a later rebuild would compute something marginally different. A number you saw here will still be that number in five years.
How much do provisional values move? Measured across the archive, a value drifts by 0.09 to 0.26 index points over its floating period — under half a percent of a reading near the baseline of 50, and smaller than the day-to-day noise in the signal itself.
The most recent days also carry a second, separate property: the current calendar month has no settled Google benchmark yet, so its days are built from the daily and weekly series alone rather than being fitted to a monthly total. They are marked provisional for that reason too, and are rebuilt onto the benchmark once the month closes.
When a signal is late
The four signals come from four independent services, and one of them can stop publishing without the others noticing. In August 2026 Wikimedia's pageview data stopped updating for three days, across every language and every project. Nothing was wrong with Wikipedia, and nothing was wrong with the other three signals — but the Index had been built to wait for all four, so the front page sat three days out of date while news, search and social had all reported.
A reader cannot tell a stale Index from a calm one by looking at it. So the rule is now:
- **The Index publishes as soon as at least three of the four signals have
reported**, rather than waiting for the fourth.
- A reading built from three signals says so, on the front page, names which
signal is missing, and says when that source last published.
- Nothing is estimated. The missing signal is left out of the average and
its weight is spread across the three that reported. It is not carried forward from yesterday, not interpolated, and not replaced with a zero. A missing measurement is recorded as missing.
- Below three signals, nothing is published. Two signals is a different
measurement wearing the same name.
- The historical series is unaffected. The chart, the downloadable data and
every comparison across time use only days on which all four signals reported, so like is always compared with like. A three-signal reading is a reading of today, not a point in the history.
The practical effect is that the headline figure is at most a day or two old essentially always, and on the rare day it is built from three signals you are told so plainly rather than being shown an old number with no explanation.
Reading the Index
- Median day: 57. Half of all days fall below this.
- Above 154: top 5% of days.
- Above 261: top 1% of days — a major event with broad public engagement.
The largest readings in the series are the Ever Given grounding in the Suez Canal (March 2021, peaking at 878) and the Baltimore bridge collapse (March 2024, peaking at 615).
Days are not weeks. 2026 is the highest-reading year in the series and March 2026 the highest-reading month, but the highest single days remain Ever Given and Baltimore. A sustained period of elevated attention and a single overwhelming shock are different things, and the Index distinguishes them.
Limitations
Stated plainly, because they affect how the numbers should be used.
The history is reconstructed, not recorded. The Index was built in 2026 and its history computed from archives — Google Trends, Wikipedia, GDELT and platform APIs all serve historical data. This is standard for an index at launch, but it is not the same as having measured each day as it happened. Readings from 21 July 2026 onward are collected daily.
Search levels are not comparable across long spans. A control basket of extremely common English words with no maritime content rises 53% over the measured period, so part of any long-run rise in the search signal reflects changes in the platform rather than in the world. Comparisons within recent years are sound; a direct 2016-to-2026 comparison of the search component is not, and roughly a third of its 2026 level is attributable to that drift.
Wikipedia's audience is changing underneath the measurement. Reference traffic is being absorbed by AI answer engines. The control basket removes most of this, since controls are subject to the same force — but "most" is not "all", and the effect is still developing.
Social is one platform in practice. X carries 90% of the component's weight. A change in its API, pricing or policy would be a material problem.
Missing data is published as missing. No signal is ever recorded as zero when it failed to report, and no value is carried forward from the previous day. A day with a failed feed publishes nothing for that component. Before 1 January 2020 the News signal does not exist at all, so the composite for 2016–2019 is published as a separate, clearly-labelled series rather than as a continuation of the same measurement.
Versioning
This is version 1.0, the first published methodology.
Every future change — a keyword, a weight, a new platform, a baseline — will be released as a numbered version, announced, and accompanied by a restatement of the affected history. Both versions will remain available. Nothing changes silently.
Data and citation
The daily series is available as JSON. Cite freely with attribution to the Maritime Attention Index and a link to this site. For component-level data or press enquiries, get in touch through the fortnightly edition.