{
  "$type": "site.standard.document",
  "bskyPostRef": {
    "cid": "bafyreifzbm52y65bhl4jmh24jwjibb6vdm4jjocuxwmwfhwbrqtdbp3osa",
    "uri": "at://did:plc:3fychdutjjusoqeq24ljch6q/app.bsky.feed.post/3mpqzfkdqcyg2"
  },
  "coverImage": {
    "$type": "blob",
    "ref": {
      "$link": "bafkreiflo6xt7is6b2iafwghkjahlgggocme5jwjsbeuqqwcywuvjhmszm"
    },
    "mimeType": "image/png",
    "size": 24783
  },
  "path": "/abs/2607.01451v1",
  "publishedAt": "2026-07-03T00:00:00.000Z",
  "site": "https://arxiv.org",
  "tags": [
    "Foad Namjoo",
    "Drew McClelland",
    "Michael Matheny",
    "Jeff M. Phillips"
  ],
  "textContent": "**Authors:** Foad Namjoo, Drew McClelland, Michael Matheny, Jeff M. Phillips\n\nAnomaly detection in geospatial data is a crucial tool in geographic information science (GIS), with applications ranging from national security to public-health surveillance to the study of societal disparities. This work focuses on spatial scan statistics and addresses a key mismatch: spatial counts are typically aggregated into predefined regions (census tracts, zip codes, counties), whereas the most efficient scan algorithms operate on spatial point data. The standard remedy -- collapsing each region to its centroid, as in widely used tools such as SaTScan -- is convenient but, as we show, discards the region's spatial extent and causes a significant loss in statistical power. To resolve this, we propose a simple yet scalable fix: replace each spatial region with 20-50 points sampled uniformly from its geometry and spread the region's values evenly across them. This approach improves statistical power while maintaining computational tractability. A convergence analysis explains why so few samples per region suffice. We recommend this sampling-based conversion as the default way to apply point-based spatial scan statistics to region-aggregated data for anomaly detection.",
  "title": "Sampling for Region-Aggregated Spatial Scan Statistics"
}