Performance appraisals are meant to be objective mirrors that reflect how well employees are doing their jobs. In the tourism industry, where service quality directly shapes guest experiences, getting these evaluations right is critical. Yet appraisals frequently fall short of their purpose due to two persistent challenges: validity and reliability. When an appraisal lacks either, it stops measuring real performance and starts generating noise that misleads managers, frustrates employees, and damages trust across the organisation.
Table of Contents
- Understanding validity and reliability in appraisals
- Validity problems that distort appraisals
- The halo and horns effect
- Personal bias and similarity effects
- Different rating patterns: leniency, strictness, and central tendency
- Recency and contrast errors
- Reliability problems that erode consistency
- Instability over time
- Inconsistencies among raters
- Lack of training in appraisal techniques
- Why these issues matter in service-driven workplaces
- Practical ways to strengthen validity and reliability
- Standardise the criteria
- Use behaviourally anchored rating scales
- Train raters and run calibration sessions
- Use multiple rating sources
- Document performance throughout the year
Understanding validity and reliability in appraisals
Before exploring the problems, it helps to clarify what these two terms actually mean. Validity refers to whether an appraisal tool actually measures what it claims to measure, while reliability is the extent to which the tool produces consistent results when used repeatedly. A valid appraisal captures the true performance of a front-office executive or a tour guide. A reliable appraisal produces similar ratings whether conducted by one supervisor or another, this month or six months from now.
An appraisal system that is high on both is the goal. Without strong validity and reliability, serious questions arise about both the usefulness and the legality of the system, especially when ratings drive decisions about promotions, increments, or termination.
Validity problems that distort appraisals
Validity issues creep in when something other than actual job performance influences the rating. In tourism workplaces, where staff are evaluated on a mix of measurable outputs and soft skills like guest interaction, these problems can quietly skew results.
The halo and horns effect
The halo effect is one of the most common rating distortions. It is the tendency to make inappropriate generalisations from one aspect of a person’s job performance, where one outstanding characteristic colours the entire evaluation. A receptionist who consistently arrives early might be rated highly on teamwork, communication, and even problem-solving, simply because punctuality has created a positive overall impression. The horns effect is the mirror image – one weak area, like slow report submission, drags down ratings on unrelated dimensions where the employee may actually excel.
Research shows the halo effect appears more often when raters lack deep job knowledge or familiarity with the employee being rated. A practical fix is to have supervisors rate different traits at separate times – for example, evaluating attendance one day and dependability another – which forces raters to consider each dimension on its own merits.
Personal bias and similarity effects
Bias enters appraisals in many forms. Supervisors sometimes allow personal preferences, dislikes, or even racial and gender biases to influence their evaluations. The similar-to-me effect is particularly subtle: managers tend to rate employees who share their background, communication style, or interests more favourably than those who do not. In a hotel where supervisors and team members come from diverse linguistic and cultural backgrounds, this bias can quietly disadvantage entire groups.
Closely related is leniency bias. According to one analysis of rater behaviour, this happens when a manager gives an inflated rating because of sympathy or empathy – perhaps knowing an employee is dealing with personal problems and not wanting to add to their stress. While the intention is kind, the result is a distorted record that makes it harder to identify real top performers.
Different rating patterns: leniency, strictness, and central tendency
Three rating patterns are commonly grouped together as distributional errors because they affect how scores spread across the rating scale.
Leniency error occurs when a rater consistently gives inflated scores. Strictness error is the opposite – every employee is rated harshly, regardless of actual performance. Central tendency error describes raters who cluster everyone in the middle of the scale. In short, the central tendency error is the failure to recognise either very good or very poor performers, and is often the default when a manager feels uncertain or wants to avoid difficult conversations.
Two raters using only narrow portions of the same scale – one harsh, one lenient – will produce wildly different ratings for the same level of work, undermining both fairness and the data the organisation relies on.
Recency and contrast errors
Performance appraisals are typically annual or biannual, but human memory is short. Recency error is the tendency to weigh recent events too heavily. A travel desk executive who handled a difficult group booking in the last week of the review cycle may be rated as a star performer, even if the previous eleven months were unremarkable. Contrast error works differently – it occurs when supervisors compare employees to one another rather than to an objective performance standard, so a competent guide working alongside an exceptional colleague may appear weaker than they actually are.
Reliability problems that erode consistency
Even when an appraisal tool measures the right things, it can fail at consistency. Reliability problems show up when the same employee receives different ratings across time periods, raters, or contexts despite no actual change in performance.
Instability over time
Ratings can drift simply because the rater’s mood, workload, or external pressures change between review cycles. A manager working through a stressful peak season may rate more harshly than the same manager during a quieter month. Appraisal reliability and validity remain major problems in most appraisal systems, and new systems are often met with substantial resistance precisely because employees sense this drift even when they cannot name it.
Time-based instability is especially relevant in tourism, where business cycles fluctuate sharply. Ratings collected during a high-pressure festival season may not align with those collected during a lean travel period, even though the underlying performance is similar.
Inconsistencies among raters
In many tourism businesses, employees report to multiple supervisors – a duty manager during the day, another at night, plus a department head. When these raters use different mental yardsticks, results diverge sharply.
The technical term for this is poor inter-rater reliability. Assessment tools that rely on ratings must exhibit good inter-rater reliability, otherwise they are not valid tests. If one supervisor weighs guest feedback heavily while another prioritises operational efficiency, the same housekeeping attendant could receive very different scores in the same review window. The appraisal stops being a measure of the employee and becomes a measure of which supervisor happened to fill in the form.
Lack of training in appraisal techniques
Many supervisors are promoted into appraisal responsibilities without ever being trained on how to conduct one. They may not recognise their own biases, may misinterpret rating anchors, or may apply standards inconsistently. One study found that raters trained using a specific methodology achieved a Cohen’s Kappa value of 0.85, indicating high agreement, compared to untrained raters at just 0.5 – a striking gap that translates directly into fairer evaluations.
Training is especially important in service industries where soft skills like empathy, communication, and cultural sensitivity matter enormously but resist easy measurement. Without preparation, raters tend to fall back on instinct, which is exactly where bias lives.
Why these issues matter in service-driven workplaces
Tourism organisations live and die by the quality of guest interaction. When appraisals fail to validly capture this performance, several downstream problems follow. Talented employees feel unrecognised and disengage. Underperformers escape detection and continue to weaken the team. Promotion and increment decisions reward the wrong people, eroding trust in management. Over time, the appraisal system becomes a ritual rather than a tool – completed for compliance but ignored for decisions.
There is also a legal dimension. When ratings influence terminations, demotions, or pay, an appraisal system that cannot be defended on grounds of validity and reliability exposes the organisation to disputes and regulatory scrutiny.
Practical ways to strengthen validity and reliability
The good news is that both problems respond well to deliberate intervention. A few strategies stand out.
Standardise the criteria
Vague performance standards invite subjective interpretation. Instead of writing ‘increase sales’ as a performance standard, a clearer version is ‘increase sales by 10 percent from last year’ – a measurable target that any rater can verify. In tourism, standardised metrics could include guest satisfaction scores, occupancy contribution, complaint resolution time, or adherence to service protocols.
Use behaviourally anchored rating scales
Generic scales of 1 to 5 leave too much room for interpretation. Behaviourally anchored rating scales, or BARS, attach specific examples of observable behaviour to each rating point. BARS define scale points with specific behaviour statements that describe varying degrees of performance, making it much clearer to both rater and employee what each score actually means.
Train raters and run calibration sessions
Bringing supervisors together to rate the same sample case and then discussing why their scores differ is one of the most powerful corrective tools. Calibration exercises engage raters in discussions to ensure consistency in ratings and scores, exposing hidden differences in interpretation before they affect real employees.
Use multiple rating sources
A single rater carries a single set of biases. Multi-source feedback – sometimes called 360-degree appraisal – combines input from supervisors, peers, subordinates, and in tourism, often guests as well. Each source corrects for the blind spots of the others, producing a more balanced and stable picture.
Document performance throughout the year
Recency and primacy errors thrive when the rater relies on memory. A simple log of significant incidents, complaints handled, commendations received, and goals met turns the annual review from an act of recall into an act of summary, which is far more reliable.
What do you think? Which of these validity or reliability problems do you suspect is the most common in tourism workplaces you have observed, and what would change in an organisation if its appraisal system suddenly became genuinely fair and consistent?
References
- https://learn.saylor.org/mod/book/view.php?id=60428&chapterid=47713
- https://opentext.ku.edu/teams/chapter/performance-evaluation/
- https://www.dartmouth.edu/hr/professional_development/for_managers/performance_management/common_rater_errors.php
- https://txwes.pressbooks.pub/iopsychologytxwes/chapter/7-3-performance-appraisal-part-2-rating-distortions/
- https://factorialhr.com/blog/bias-in-performance-reviews/
- https://bizfluent.com/about-5445066-importance-reliability-performance-appraisals.html
- https://en.wikipedia.org/wiki/Inter-rater_reliability
- https://encord.com/blog/inter-rater-reliability/
- https://www.nationalforum.com/Electronic%20Journal%20Volumes/Lunenburg,%20Fred%20C.%20Performance%20Appraisal-Methods%20And%20Rating%20Errors%20IJSAID%20V14%20N1%202012.pdf
- https://www.numberanalytics.com/blog/ultimate-guide-to-inter-rater-reliability
Leave a Reply