Government Web Relationship Data
This page provides comprehensive insights into the web relationships between European and allied government websites. We analyze links, scripts, styles, forms, and media dependencies to understand how government sites connect and interact with each other and third-party services.
What is Extracted
This dataset contains relationship observations from government web pages scanned as part of the EU Government Accessibility Statement Discovery project. Relationships include:
- Editorial Links:
<a href>links to internal/external pages - Script Dependencies:
<script src>external JavaScript resources - Style Dependencies:
<link rel="stylesheet">CSS stylesheets - Font/Prelooad Dependencies:
<link rel="preload"|"font"|"preconnect">font resources - Image/Media Dependencies:
<img src>images and media - Form Destinations:
<form action>external form submissions
Classification Categories
External domains are automatically categorized based on their function:
- Analytics: Google Analytics, Tag Manager, Heatmaps, etc.
- Social Platform: Facebook, Twitter/X, LinkedIn, Instagram, YouTube
- CDN: Content Delivery Networks like Cloudflare, Fastly, Akamai
- Identity: Authentication services like Okta, Auth0, Microsoft Login
- Document Host: File hosting services like Google Drive, Dropbox, SharePoint
- Commercial Service: Payment, support, and business services
Methodology Notes
- Graph Metrics: In-degree, out-degree, and weighted measures represent observed web relationships, NOT institutional authority or hierarchy.
- Country Assignment: Prioritized through institutional evidence, not just TLD mapping.
- Classification Precedence: Authoritative registry > Curated source > Institutional source > Structured source > Software Heritage > Domain pattern > Network signal
- Data Provenance: All data is sourced from machine-readable scan results with full tracking of observation timestamps and source pages.
Filters and Search
Total Domains
Total Relationships
Countries
Unique Targets
Editorial Links
Technical Dependencies
| Country | Domain | Targets | Top Types | Observations | Last Seen |
|---|---|---|---|---|---|
No domains match the current filtersTry adjusting your search criteria or filters |
|||||
| Source Domain | Target Domain | Relationship Type | Source Pages | Observations | Page Regions | First Seen | Last Seen |
|---|---|---|---|---|---|---|---|
No relationships match the current filtersTry adjusting your search criteria or filters |
|||||||
Loading country data...
By default this shows government-to-government editorial links only for the selected country โ no analytics, CDNs, or other third-party services. Pick a country above and select "Include third-party services" to see the rest.
Graph Notes: Node size = connection count. Edge thickness = observation count. This visualization represents observed web relationships and does NOT establish institutional authority or government hierarchy.
Downloads
Download the complete relationship dataset or specific filtered results for offline analysis.
Data Extraction Process
All relationship data is extracted during the scanning process using the existing MultiScanner infrastructure. The HTML parsing is performed using BeautifulSoup and tldextract for domain normalization. URLs are resolved relative to their source page and normalized to eliminate fragments, default ports, and IDNA encoding.
Classification Logic
Domain categorization follows a configurable rule-based approach defined in data/relationship_categories.json. Categories are applied to registrable domains and represent preliminary indicators rather than definitive classifications.
Aggregation Methodology
For each scan batch, relationships are deduplicated within source pages and then aggregated by:
- Source registrable domain
- Target registrable domain
- Relationship type
Limitations and Interpretations
- Graph metrics (in-degree, out-degree) indicate observed connection patterns, not institutional authority or hierarchy.
- Domain-level analysis does not infer central government status from network structure.
- Country assignment prioritizes institutional evidence over TLD mapping alone.
- Classification uses preliminary categorization that may require manual verification for certain services.
- Data represents a snapshot in time; relationships may change as pages are updated.
Provenanced Data
All data maintains full provenance tracking with:
- Source page URLs
- Extraction timestamps
- Page region information
- Canonical domain records with classification basis
- Multiple source evidence for each domain
Submitting Corrections
If you identify errors in the relationship data, please use the issue tracker on GitHub to submit corrections. Include the source page URL, target domain, and the specific type of correction needed (classification, country assignment, relationship type, etc.).
Data Schema Version: 1.0.0
Last Updated: 2026-07-15