AI travel agents have a data problem before they have an AI problem
Language models understand the traveler's request. The harder problem is providing trustworthy product identity, pricing, availability, provenance and freshness.
Field notes on keeping data pipelines alive, getting AI agents past the demo, and building software that survives an audit. No thought leadership — just what we ran into and what it cost.
Language models understand the traveler's request. The harder problem is providing trustworthy product identity, pricing, availability, provenance and freshness.
The useful question in aviation data is not API versus scraping. It is which source should be authoritative for each requirement — and how to normalize the result.
The same experience can appear as several listings, options, languages and price structures across different systems. Collecting the pages is only the beginning.
A collector proves that data can be extracted. A production travel-data pipeline has to keep the dataset complete and trustworthy while the systems underneath it keep changing.
The most dangerous external-data failure is the one that looks successful: the job ran, requests returned 200, rows landed — and the dataset is wrong.
Most scraping projects work perfectly for about eight weeks. What happens after that is predictable, and it has almost nothing to do with how well the scraper was written.
AI pilots stall for reasons that have very little to do with which model was chosen. The failure is almost always in the layer underneath — and it was there before the pilot started.
Access control and audit trails are the least demonstrable parts of regulated software — and close to impossible to retrofit convincingly. A note on why they go in first.
If a pipeline keeps breaking or an AI pilot won't ship, we've probably seen that specific version of it. Thirty minutes, no deck.
Talk to an engineer → RSS