AI for Travel & Hospitality

AI for High-Volume Travel and Hospitality

Insights from Fin Team•

AI for high-volume travel and hospitality - title image

Your queue is mostly the same handful of questions. Where's my booking. Cancel this. Change my dates. Why hasn't my refund landed. The volume climbs every time the business grows, and three or four times a year it climbs several times higher than that, in a window you cannot staff for without carrying idle headcount the rest of the year.

If you have already tried an AI tool for this, you probably know how that went. Most travel and hospitality teams evaluating AI in 2026 are on their second or third attempt. The last one deflected tickets, could not reach a live booking record, and resolved almost nothing.

So the useful question is: what has to be true for it to work this time?

This guide covers that: what separates an agent that resolves from one that deflects, how to evaluate against the failure mode that burned you before, and what the economics look like against what a human contact costs you today.

Key takeaways:

  • The difference between deflection and resolution is live system access. An agent that cannot read your booking record and write back to it can only answer questions.
  • Peak season and disruption events are the real test. Elastic capacity during a holiday rush or an IROPs day is worth more than any headline resolution percentage.
  • Hard policy controls matter as much as accuracy. Most teams' sharpest fear is a wrong answer screenshotted onto social.
  • The business case is headcount dependency, not efficiency. Cost per resolution has to beat the fully loaded cost of a human handling that contact.
  • Start narrow. One high-volume query type, proven on your real data, then expand.
  • Spend is already committed sector-wide. The AI in hospitality and tourism market sits at $26.53 billion in 2026, forecast to reach $75.66 billion by 2030 at a 29.9% CAGR. The question is no longer whether to invest but whether your deployment resolves anything.
Hospitality & Tourism market set to triple by 2030

Why high-volume travel support breaks at peak

Staffing support is not a headcount decision. It is a forecast with too many unknowns. Demand is hard to predict. Seasonality creates spikes that are difficult to size in advance. Attrition is a variable. And hiring has to happen before the need arrives, because interviews, onboarding, and ramp all take time.

Get any part of that wrong and there is no cushion. Over-hire and margins disappear into idle headcount. Under-hire and the queue collapses in the week that matters most.

Peak season, IROPs, cancellations, and refund spikes

Travel has two kinds of volume spike, and they break a team in different ways.

The first is seasonal and forecastable in shape if not in size: school holidays, summer, Black Friday for OTAs and package operators, the run-up to a major event. You know roughly when it is coming. You do not know how big, and you have to commit to headcount months before you find out.

The second kind arrives without warning. An IROPs event, in airline language: irregular operations. Weather closes a hub, a crew shortage cascades, a systems outage grounds a fleet. Thousands of itineraries break simultaneously, and every affected passenger contacts you inside the same hour. Volume does not rise by 30%. It multiplies, in minutes, and every contact is urgent, emotional, and specific to one booking.

No staffing model absorbs that. The math does not work at any headcount you could justify carrying through a normal Tuesday. What follows is a queue measured in days, a wave of one-star reviews, and refund and rebooking requests that keep compounding while the backlog clears.

Cancellations and refunds have their own pattern. They cluster after disruption, after policy changes, and after any pricing or fee announcement that makes customers reconsider. They are also the contacts most likely to escalate if handled slowly, because money is already out the door.

Cost per contact vs resolution rate

Two numbers get confused here, and the confusion is expensive.

Resolution rate tells you what proportion of conversations get fully solved without a human. Cost per contact tells you what each one costs you. A high resolution rate on a tool priced per conversation can still cost more than the queue it replaced. A low resolution rate at any price means the volume comes back to your team anyway, having burned a step first.

The number that decides anything is cost per resolved contact against the fully loaded cost of a human handling that same contact. And "fully loaded" is where most teams undercount, because the wage is the smallest part of the total.

The queue is also repetitive by design, which shapes who you can hire for it. Where's my booking, cancel this, change my dates. Because the work is simple, teams get built accordingly: offshore, recent grads, people between other roles. Because the work offers little development, people leave constantly. Every departure means re-hiring, re-onboarding, and re-ramping, which eats margin you did not have to spare and makes the capacity forecast worse. A team cycling through new hires is slower to perform and harder to plan around, and that cost belongs in your per-contact number too.

How B2C support earns customer trust in AI

The 2026 AI Sentiment Report

We surveyed 1,026 end users to understand how they feel about AI in customer service. Positive sentiment rose once they saw an Agent resolve a real query.

Why your last AI attempt failed

The failure is almost always the same, and it is not the model.

It could not reach live data. A decision-tree or FAQ bot sitting on top of a help center has no idea which booking the guest is asking about. It can recite your cancellation policy. It cannot look up the reservation, apply the policy to it, and process the refund. That gap is the entire difference between deflection and resolution, and it is why the last tool's numbers looked fine while your queue did not move.

The knowledge base was not ready. An agent resolves based on the content you give it. Outdated, thin, or contradictory content produces outdated, thin, and contradictory answers. This is worth being blunt about: no AI tool will pass an accuracy test on a neglected help center. Content quality is the single biggest lever on performance, and it is the one most teams discover late.

Nobody owned it after launch. The vendor implemented, handed off, and disappeared. Performance plateaued because no one was analyzing conversations, finding gaps, updating content, and testing changes on a regular cycle.

Industry data shows how widespread this is. Wyndham's 2026 Owner Trends Report found 98% of hotel owners have begun incorporating AI, but only 32% have it embedded across most of their operations, and 73% want to do more while feeling overwhelmed about where to start. BCG's AI-First Hotels analysis with NYU puts fewer than 10% of hospitality companies in the tier where AI generates substantial value, with roughly 25% at the scaling stage below it.

Adoption without resolution. The tools were sold faster than they could deliver, and teams running high-volume queues absorbed most of that cost.

The spending is not slowing down. Research and Markets puts the AI in hospitality and tourism market at $26.53 billion in 2026, up from $20.39 billion the year before, and forecasts $75.66 billion by 2030 at a 29.9% CAGR. Customer service and support is one of the largest application segments in that spend. Which means the risk for a CX leader is no longer being left behind. It is being the person who bought the second tool that did not work.

Why your last customer service AI Agent attempt did not move the queue

What AI customer service means for airlines, OTAs, and economy hospitality

AI customer service agent vs chatbot FAQ deflection

Deflecting vs Resolving

These are different categories of software that get sold with similar language.

A deflection chatbot matches a question to an article and shows it. Success is defined as the customer not opening a ticket. It has no idea who the customer is, what they booked, or whether their problem got solved. The metric it optimizes, deflection, counts a customer who gave up as a win.

An AI agent takes an action. It authenticates the passenger or guest, reads the specific reservation, applies your policy to that record, executes the change in your booking system, and confirms it. Success is a resolved contact.

The practical test is one question: can it cancel a real booking and refund it to the original payment method? If the answer involves showing the customer a policy article, you are looking at a deflection tool with a language model attached.

This distinction matters more in travel than in most categories, because almost nothing in a travel queue is a general question. "What is your change fee" is rare. "What is the change fee on my flight, given the fare class I booked and the fact that I am flying in nine hours" is the actual contact volume.

Chat, voice, and after-hours coverage without linear headcount

Travel does not keep business hours. A cancelled connection at 2am generates the same urgency as one at 2pm, and a guest three time zones away does not know or care where your support team sleeps.

Covering that with people means either a follow-the-sun operation, which is expensive and hard to staff consistently, or an offshore night shift handling your most stressed customers with the least context. Most high-volume teams do the second and know it is a weak point.

An AI agent breaks the link between coverage and headcount. The same policies apply at 3am as at 3pm, and after-hours volume gets resolved rather than queued for the morning shift to inherit.

Channel strategy should follow the same logic. Chat and email are the baseline. WhatsApp matters for international travel and is the default messaging channel across much of Latin America, Europe, and Asia. SMS matters for disruption notifications. Voice remains a primary channel in travel, especially during disruption, and it is worth planning for. That said, most teams do better proving the agent on chat and email first, then extending to voice once it has earned trust, since voice economics and failure modes differ.

The requirement is that the platform extends to a new channel without adding a vendor to the stack.

How B2C support earns customer trust in AI

The 2026 AI Sentiment Report

We surveyed 1,026 end users. Positive sentiment rose once they saw an AI agent resolve a real query.

High-volume use cases that move the needle

Booking changes, cancellations, and refunds

The highest-volume, most automatable category in travel, and the one where retrieval-only tools fail most visibly.

Resolving these end to end means authenticating the customer, pulling the specific booking, evaluating it against the applicable policy including fare rules or rate plan conditions, calculating any fare difference or penalty, executing the change in the booking system, processing the refund to the original payment method, and confirming in writing.

Every step needs live system access, and several need hard limits. A refund above a threshold should require approval rather than proceeding because the conversation seemed to warrant it.

Delays, rebooking, and disruption (IROPs) at scale

The case that justifies the whole investment, because it is the scenario no staffing model covers.

During an IROPs event the agent should identify affected bookings, explain what happened and what the customer is entitled to under your policy and the applicable regulation, offer available rebooking options against live inventory, execute the rebooking, and handle downstream consequences: connecting segments, hotel nights, seat reassignment, refund of unused portions.

Two things make this work. Live inventory access, so options offered are options that exist. And elastic capacity, so the agent still performs when volume multiplies rather than degrading exactly when it is needed.

Passenger rights regimes such as EU261 and equivalent rules elsewhere make policy accuracy legally material here, not just a service concern. This is a place where an agent's answer can create a liability, which is why the guardrail question and the audit trail question matter.

Multilingual passenger and guest support without multiplying agents

A disruption at one international hub generates simultaneous contacts in a dozen languages. Staffing for that means hiring per language and per shift, and most teams cover a handful of languages well and the rest through translation tooling that adds delay and loses nuance.

What you want is automatic language detection with native-quality response, from one deployment rather than a separate bot per market. The test is whether a Portuguese-speaking customer gets the same policy accuracy and the same ability to take action as an English-speaking one, not just a translated apology.

Contact center AI and customer service automation for travel ops

The agent is one component. How it sits in your operation determines whether it holds.

Deflection that still resolves

The reason "deflection" carries a bad reputation in travel is that most of it was dead-end deflection: the customer was routed away from a human without getting an answer, so they came back angrier, or left and posted about it.

Reducing human contacts is a legitimate goal. The distinction is whether the contact ends resolved or abandoned. Measure genuine resolution, and watch repeat contact rate alongside it. A resolution rate that rises while repeat contacts also rise is not resolution, it is a customer trying again.

Design for the handoff too. When the agent cannot or should not resolve something, the handoff should carry full conversation history, the booking context, and a summary, so the customer never restarts. Fraud claims, safety issues, bereavement cases, and anything with legal exposure should route to a human by rule rather than by the agent's judgment.

Omnichannel coverage and the ops stack

Practical questions to settle before deployment.

Where does the agent sit relative to your helpdesk? Running an AI layer on top of an incumbent platform means two contracts and an integration to maintain. Consolidating onto one platform removes that, and it is usually where the cost case becomes decisive.

How does it affect workforce planning? If the agent handles a predictable share of a query type, your peak staffing model changes. Model that explicitly rather than treating automation as headroom.

Who owns performance after launch? Someone has to review conversations, find content gaps, test changes, and redeploy on a regular cycle. Name that person before you sign, not after.

What does the audit trail look like? For refunds, policy decisions, and anything touching passenger rights, you need to reconstruct why the agent answered as it did.

Budget airlines, OTAs, and high-volume booking platforms: what "good" looks like

Volume pressure by business model

The shape of the problem varies by model, though the underlying pattern holds.

Low-cost carriers run high passenger volume on thin per-ticket margins, with ancillary revenue doing much of the work. Support volume concentrates in bag fees, seat selection, change fees, and disruption, and the per-contact economics are unforgiving: a handled contact can cost a meaningful share of the fare.

OTAs and metasearch platforms carry a structural complication. A single itinerary can involve an airline, a hotel, a car rental, and an activity provider, each with its own policy, API, and refund process. Resolving one cancellation may mean reconciling four suppliers, and the customer holds the OTA responsible for all of them.

Economy and midscale accommodation groups run high booking volume across many properties, often with lean central support and property staff who cannot absorb overflow. Volume clusters around booking modifications, cancellation policy, payment questions, and arrival logistics.

Tour operators, activities, and ticketing face weather-driven cancellation waves and tiered refund policies that are hard for staff to apply consistently, let alone quickly.

Volume pressure by business model

Metrics: containment, CSAT under load, cost to serve

Four numbers, read together.

  1. Containment or resolution rate, defined as genuinely resolved rather than merely not escalated. Ask any vendor precisely how they define a resolution before comparing numbers, because definitions vary and the differences are large.
  2. CSAT under load, not CSAT on average. Segment satisfaction for your peak period and for disruption contacts specifically. An agent that scores well in a normal week and collapses during an IROPs event has not solved your problem.
  3. Cost to serve per resolved contact, against the fully loaded human comparison below.
  4. Repeat contact rate, as the check on whether resolution is real.

Set a baseline on all four before you pilot. Without a pre-deployment number, you will not be able to prove the result internally, and this is a buying committee that wants proof.

DEMO

Watch Fin cancel a real booking

See it authenticate the guest, read the reservation, apply your policy, and refund the original payment method.

How to evaluate an AI customer service agent for travel and hospitality

Policy-accurate resolution, handoff, and auditability

Accuracy and control are different problems. An agent can be accurate on average and still produce the one answer that ends up on social media, or the one refund commitment you now have to honor.

What you need are hard constraints rather than guidance: never refund above a set amount without approval, never contradict the published fare rules or cancellation policy, never discuss a competitor, always escalate fraud, safety, and bereavement to a human. These have to be enforced outside the model, not instructions it weighs against everything else in a persistent conversation.

Then test that they hold. Ask the vendor to demonstrate what happens when the agent is pushed: a customer insisting the policy is different, a request just above your threshold, a claim designed to trigger a payout.

Auditability is the other half. When an answer is wrong you need to see which content produced it and which policy path the agent followed, so you can fix the cause. For passenger rights and refund decisions, that record may also be what you show a regulator.

Integration with booking, PSS, and helpdesk systems

Travel POC Scorecard

Test this first, because it is where the last attempt broke.

Name your systems and make the vendor answer specifically. Airlines and OTAs: your passenger service system or GDS, whether that is Amadeus, Sabre, Travelport, or something proprietary, plus payment and loyalty. Accommodation: your PMS, whether Opera, Mews, Cloudbeds, or another, plus booking engine and channel manager. Everyone: your helpdesk and your payment gateway.

Two questions separate real integration from a roadmap promise. Can it read a specific record, not just static content? And can it write back, executing a change rather than describing one? If either answer requires weeks of custom engineering, you have learned what you need to know.

Use this as the scorecard when you run a POC:

CriterionWhat to testPassFail
Integration and actionAsk the agent to cancel a real booking and refund it against the original payment methodReads the record and writes back to it via a standard connector, configured in daysNeeds custom engineering to read a reservation, or can only quote the policy
Accuracy on your dataRun it against your live help center and 100 messy historical contactsResolves your real contacts, and the vendor is specific about which content gaps hurt performanceOnly performs on curated demo content, or claims content quality is irrelevant
Control and guardrailsPush it outside policy: a refund above threshold, a competitor question, a fraud claim, a customer insisting the rules differEnforced limits and approval thresholds hold, restricted topics escalate every timeLimits are prompt instructions that persistence can talk it past
Load handlingAsk what happens at five times normal volume, and request a reference who ran a comparable peak or IROPs eventElastic scaling with stable latency and accuracy under load, plus a customer who survived their own peakUptime figure on a slide and no peak-season reference
Handoff and auditabilityEscalate mid-conversation and inspect what the human receives; then trace why a given answer was returnedFull history, booking context, and summary transfer; answer traceable to source contentHuman restarts from scratch; no record of which content drove the answer
Channel coverageMap your current channels plus the next one you expect to addCovers today's channels and extends without a new vendor or contractChat-only, or each new channel is a separate product

Cost sits underneath all of it as a gate rather than a decision criterion. Priced too high or too unpredictably and you never reach the real evaluation. But a cheaper agent that cannot take action is not cheaper.

The business case: cost per resolution against a human contact

Most teams do not build a formal model up front. They reason backwards from people: fewer seasonal hires, a smaller offshore team, not needing the next headcount add to clear peak.

Work it out with your own numbers:

Cost inputYour number
Base wage or BPO rate per agent (annual)
Recruiting and hiring cost per agent
Onboarding and training
Ramp cost (partial productivity during weeks 1 to 8)
Management and QA overhead allocated per agent
Tooling and helpdesk seat licenses
Attrition replacement (turnover rate x cost to replace)
A. Fully loaded annual cost per agent
B. Contacts handled per agent per year
Fully loaded cost per contact (A divided by B)

Then set that figure against per-resolution pricing at your real volumes, peak included. Most teams find the gap is wider than they assumed, because the wage line is the smallest part of the total.

The second half of the case is consolidation. Running an AI agent alongside an incumbent helpdesk means paying for both, plus the integration friction between them. Replacing the helpdesk instead of adding to it collapses two contracts into one, which is often what makes the numbers work in a thin-margin business even when a single line item looks higher.

Both halves point the same direction. The case that closes is not "our accuracy is a few points better." It is "this gets us through peak without hiring a seasonal team."

TALK TO SALES

Go through peak without a seasonal team

We will model your cost per resolved contact against what a human contact costs you today, at your real volumes.

10 reasons Fin fits this profile

  1. It executes on live booking data. Data Connectors give Fin access to your booking engine, PMS, order management, payment, and CRM systems, so it can read a specific record and act on it. Procedures handle the multi-step work: authenticate the customer, retrieve the booking, apply your cancellation policy or fare rules, process the refund, confirm. Procedures combine natural language instructions with deterministic controls for the steps that need exact logic, which is how a policy becomes a constraint rather than a suggestion.
  2. It performs on real content. Fin runs on Fin Apex, a proprietary model built for customer service rather than adapted from a general-purpose one, and trains on your existing knowledge base. The Fin AI Engine retrieves, ranks, and validates each answer against your content before it returns. Aftersell's VP of FP&A, Guang Li: "What was most impressive for me is that from the moment we turned it on, we were able to achieve something like a 40% resolution rate just with our existing documentation."
  3. It holds through peak. 99.97% uptime with real-time elastic scaling. Jukebox had Fin handle 90% of peak season queries. MPB's Head of Global CX, Chris Beattie, on their busiest quarter: "Last quarter was our busiest yet, with a peak of over 40,000 conversations a month. With more and more customers sending us gear, it can put significant pressure on our operational teams. Fin really helped lighten the load."
  4. It stays inside your policies, and shows its work. Procedures enforce approval thresholds and policy logic on financial actions. Guidance sets topic restrictions, tone, and vocabulary rules across every conversation. Monitors score every conversation against your own scorecards, so performance stays visible rather than drifting quietly, and every conversation is logged for audit.
  5. It hands off properly. When Fin escalates, it passes full conversation history, an AI-generated summary, and customer context. The passenger or guest does not restart, which matters most in exactly the conversations you route to a human: disruption, fraud, bereavement.
  6. It covers your channels. Live chat, email, WhatsApp, SMS, social messaging, and voice, with context carrying across channels. Most teams start on chat and email and extend as trust builds. Fin also supports over 45 languages with automatic detection, so a disruption generating contacts in a dozen languages does not need a dozen deployments.
  7. Your team can run it. Fin is configured and maintained without engineering resources, which matters when there is no dedicated technical team to babysit a platform. Simulations let you test specific scenarios, from a weather cancellation to a disputed charge to a customer pushing for an out-of-policy refund, before anything reaches a guest.
  8. One platform instead of two. Fin is the only AI agent with a native helpdesk, so AI and human agents work in the same system rather than across two contracts and an integration.
  9. Pricing you can model. $0.99 per outcome. You pay when Fin resolves, not for conversations it does not resolve, and spend caps prevent surprises during peak. Costs track resolution volume rather than seats, which suits a seasonal business.
  10. Someone is still there after launch. Performance comes from a cycle of analyzing conversations, updating content, testing, and redeploying. Fin's team stays involved through that, which is the part most teams say their previous vendor skipped.

Fin averages a 76% resolution rate across 12,000+ customers, with many above 85%, and the average improves roughly a point a month. It holds SOC 2, ISO 27001, and ISO/IEC 42001 certifications.

Nuuly's Senior Director of Customer Success, Natalie Hurst: "We've seen a 10% increase in Fin resolution rate, which equates to about 20,000 conversations on a monthly basis, and we've also seen a 30% increase in Fin CSAT, which is really mind blowing."

Getting started

Pick one query type. Your highest-volume, most policy-defined category. For most travel and hospitality teams that is booking status, cancellations, or refunds. Do not switch everything on at once.

Check that the data is reachable. Confirm which systems hold the records that query type needs, and whether the agent can read and write to them. If it cannot, nothing else matters.

Baseline your four metrics. Containment, CSAT, cost per resolved contact, repeat contact rate. Measured before you deploy, so the result is provable afterwards.

Audit the content behind that one query type. Not the whole help center. Just the policies and articles the pilot depends on. This is the fastest version of the content work, and it tells you what the full job looks like.

Test on your real conversations. Run the agent against actual historical contacts and your live records. Measure genuine resolution, not deflection.

Set the guardrails before you go live. Approval thresholds on refunds, topic restrictions, mandatory escalation for fraud, safety, and bereavement. Test the edge cases in simulation.

Expand from a proven win. Once one query type holds up through a normal week and ideally a busy one, add the next.

FAQ

How can AI help airlines?

The highest-value applications are disruption handling and self-service booking changes. During an IROPs event, an AI agent can identify affected bookings, explain entitlements under your policy and the applicable passenger rights regulation, offer rebooking options against live inventory, and execute the change, at a volume no staffing model absorbs. Outside disruption, the volume sits in change and cancellation requests, bag and seat questions tied to a specific fare, refund status, and check-in problems. The requirement in every case is integration with the passenger service system or GDS, plus enforced policy controls, since fare rules and passenger rights carry legal consequences when answered wrongly.

Is AI replacing customer service in travel?

It is changing what human agents work on rather than removing them. Repetitive, policy-defined, record-specific contacts are what AI resolves well, and those make up the bulk of a high-volume travel queue. What stays human is the work that needs judgment or care: complex multi-supplier problems, bereavement and medical cases, fraud, complaints with legal exposure, and high-value customer relationships. In practice most teams reduce dependency on seasonal and offshore hiring rather than cutting their core team, and the people who remain handle more interesting work, which helps with the turnover problem that plagues these queues.

Do OTAs use AI for customer service?

Yes, and OTAs face a harder version of the problem than most. One itinerary can span an airline, a hotel, a car rental, and an activity provider, each with separate policies, APIs, and refund timelines, while the customer holds the OTA accountable for the whole trip. That makes multi-supplier reconciliation the main integration challenge and the main source of contact volume. It also raises the stakes on policy accuracy, since the agent has to apply the correct supplier's rules to the correct segment.

Can an AI agent actually process a refund or rebooking, or does it just answer questions?

It depends entirely on system access. An agent connected to your booking engine and payment system through data connectors can look up the specific record, apply your policy, and execute the action. An agent working only from help center content can describe a policy but cannot act on it. This is the single most important thing to test, and it is where most previous-generation tools failed.

Our last AI tool failed. What would be different?

Ask two questions of any vendor. First, can it read and write to the systems holding live booking and payment data, demonstrated on your stack rather than described. Second, what does the post-launch cycle look like and who runs it. Most failures trace to one of those two gaps rather than to model quality.

What happens during a peak event or a mass disruption?

Ask what happens to latency and accuracy at several times normal volume, and ask for a reference customer who ran the platform through a comparable spike. Reference calls are worth more here than an uptime figure. Fin maintains 99.97% uptime with real-time elastic scaling, and Jukebox handled 90% of peak season queries with it.

How do we stop the agent from giving an out-of-policy answer?

Through enforced controls rather than instructions: approval thresholds on financial actions, topic restrictions applying to every conversation, mandatory escalation for fraud, safety, and bereavement, and pre-launch simulation against edge cases. Ask to see what happens when the agent is pushed outside its policy, not only what happens when it works.

Do we have to replace our current helpdesk?

No. Fin works alongside existing helpdesks including Salesforce, Freshdesk, and Hubspot. But if you are running an AI layer on top of an incumbent platform, you are paying for both and maintaining the integration between them. Consolidating onto one platform is usually where the cost case gets decisive, and the natural time to look at it is your renewal.

How B2C support earns customer trust in AI

The 2026 AI Sentiment Report

We surveyed 1,026 end users to understand how they feel about AI in customer service. Positive sentiment rose once they saw an Agent resolve a real query.

Related articles