Understanding how real-world information sources behave.
Routing real-world information requires empirical measurement. PlaceRouter runs experiments across domains, tests routing strategies against observed outcomes, and reports the result whether or not it supports the thesis.
Replaying transit information with controlled delays, error rose as the information aged. Freshness has a measurable information cost.
Method
Routing strategies are scored against observed outcomes.
Each experiment defines a field, a set of locations, candidate sources and an observed ground truth. Confidence intervals are reported. Inconclusive results are labeled inconclusive.
Image · NASA
Weather · multi-source routing and local behaviorActive research
Local information helps. Naive per-station transfer can hurt.
20stations
580observations
24 hcheckpoint
Leave-one-station-out, mean absolute error by strategy (lower is better)
Full station lookup1.1103
Local 12h1.1864
Regional1.2542
Global1.3142
Nearest1.3699
Climate / elevation1.438
20 stations · each held out in turn
24-hour checkpoint
PlaceRouter1.2254
Best global fixed source1.2632
Finding. Local 12-hour and full station lookup outperformed the global source in the leave-one-station-out test, but climate/elevation and nearest-station transfer were worse than global. Personalization is not automatically better.
Caveat. At the 24-hour checkpoint PlaceRouter's 1.2254 against the best global fixed source's 1.2632 has a confidence interval that includes zero. Suggestive, not conclusive. We do not claim PlaceRouter beat the global source.
Air quality · PM2.5 routing, correction and calibration transferActive research
Naive routing did not beat the raw best source. Hardware-class calibration transferred.
45stations
10,459observations
117 hwindow
45 stations · best raw source at each
Action ladder, mean absolute error (lower is better)
Constrained oracle0.328
Raw best source1.284
Route + correct1.356
Correct only1.366
Finding. One source (CAMS) was best at 45 of 45 tested stations, with 80.1% dominance. Routing plus correction (1.356) and correction alone (1.366) were both worse than the raw best source (1.284). A constrained oracle at 0.328 shows the headroom (0.371) that exists but was not captured.
Calibration transfer. Calibration learned per hardware class (for example sc_sds011, sc_sps30) transferred to sensors not seen in training. Negative transfer occurred in 6 of 43 cases.
Caveat. This is a single dataset and window. It argues against naive correction, and for learning behavior per hardware model rather than per station.
Transit · controlled freshness replayActive research
Freshness has a measurable information cost.
15,318paired observations
MAE ↑as information aged · chart above
Finding. Replaying transit information with controlled delays, error worsened as the information became stale. Finding a source that contains a field is not enough; the age of the information matters.
Caveat. The effect size depends on the feed and the field. The direction was consistent in this replay.
Multi-domain source probeCompleted probe
Access, authentication, stability and authority differ substantially by domain.
32sources
17verticals
32 sources · observed status
15 open, no credential4 free key8 unreachable or shape changed5 private
Finding. Promising open or first-party categories included marine, flood and water, road traffic, transit and EV charging. A quarter of tested sources were unreachable or had changed shape, which is itself a routing problem.
Hospitality · field-level availability auditActive research
Structured availability varies dramatically by field.
44hotels
110pages
Share of hotels where each field was machine-readable · human-readable only · not found
Address100% · 0% · 0%
Phone90.9% · 9.1% · 0%
Check-in9.1% · 9.1% · 81.8%
Check-out9.1% · 11.4% · 79.5%
Pets2.3% · 18.2% · 79.5%
Parking fee0% · 29.5% · 70.5%
Breakfast0% · 31.8% · 68.2%
machinehuman onlynot found
Finding. Address and phone are almost always machine-readable. Policy fields such as check-in, parking fees, breakfast and pets are mostly absent or human-only. Field-level routing has to expect this unevenness in every domain.
The router has to learn per domain.
Per-hardware calibration was useful in air quality. Per-station personalization was harmful in weather. Freshness cost accuracy in transit. These are exactly the kinds of domain-specific source behaviors PlaceRouter is designed to learn rather than assume.
AI can reason. PlaceRouter connects it to the world.
Have a dataset, a sensor network, or an information requirement worth measuring? Research collaborations and pilots start the same way.