Performance appraisals are meant to be objective mirrors that reflect how well employees are doing their jobs. In the tourism industry, where service quality directly shapes guest experiences, getting these evaluations right is critical. Yet appraisals frequently fall short of their purpose due to two persistent challenges: validity and reliability. When an appraisal lacks either, it stops measuring real performance and starts generating noise that misleads managers, frustrates employees, and damages trust across the organisation.

Table of Contents

Understanding validity and reliability in appraisals

Before exploring the problems, it helps to clarify what these two terms actually mean. Validity refers to whether an appraisal tool actually measures what it claims to measure, while reliability is the extent to which the tool produces consistent results when used repeatedly. A valid appraisal captures the true performance of a front-office executive or a tour guide. A reliable appraisal produces similar ratings whether conducted by one supervisor or another, this month or six months from now.

An appraisal system that is high on both is the goal. Without strong validity and reliability, serious questions arise about both the usefulness and the legality of the system, especially when ratings drive decisions about promotions, increments, or termination.

Validity problems that distort appraisals

Validity issues creep in when something other than actual job performance influences the rating. In tourism workplaces, where staff are evaluated on a mix of measurable outputs and soft skills like guest interaction, these problems can quietly skew results.

The halo and horns effect

The halo effect is one of the most common rating distortions. It is the tendency to make inappropriate generalisations from one aspect of a person’s job performance, where one outstanding characteristic colours the entire evaluation. A receptionist who consistently arrives early might be rated highly on teamwork, communication, and even problem-solving, simply because punctuality has created a positive overall impression. The horns effect is the mirror image – one weak area, like slow report submission, drags down ratings on unrelated dimensions where the employee may actually excel.

Research shows the halo effect appears more often when raters lack deep job knowledge or familiarity with the employee being rated. A practical fix is to have supervisors rate different traits at separate times – for example, evaluating attendance one day and dependability another – which forces raters to consider each dimension on its own merits.

Personal bias and similarity effects

Bias enters appraisals in many forms. Supervisors sometimes allow personal preferences, dislikes, or even racial and gender biases to influence their evaluations. The similar-to-me effect is particularly subtle: managers tend to rate employees who share their background, communication style, or interests more favourably than those who do not. In a hotel where supervisors and team members come from diverse linguistic and cultural backgrounds, this bias can quietly disadvantage entire groups.

Closely related is leniency bias. According to one analysis of rater behaviour, this happens when a manager gives an inflated rating because of sympathy or empathy – perhaps knowing an employee is dealing with personal problems and not wanting to add to their stress. While the intention is kind, the result is a distorted record that makes it harder to identify real top performers.

Different rating patterns: leniency, strictness, and central tendency

Three rating patterns are commonly grouped together as distributional errors because they affect how scores spread across the rating scale.

Leniency error occurs when a rater consistently gives inflated scores. Strictness error is the opposite – every employee is rated harshly, regardless of actual performance. Central tendency error describes raters who cluster everyone in the middle of the scale. In short, the central tendency error is the failure to recognise either very good or very poor performers, and is often the default when a manager feels uncertain or wants to avoid difficult conversations.

Two raters using only narrow portions of the same scale – one harsh, one lenient – will produce wildly different ratings for the same level of work, undermining both fairness and the data the organisation relies on.

Recency and contrast errors

Performance appraisals are typically annual or biannual, but human memory is short. Recency error is the tendency to weigh recent events too heavily. A travel desk executive who handled a difficult group booking in the last week of the review cycle may be rated as a star performer, even if the previous eleven months were unremarkable. Contrast error works differently – it occurs when supervisors compare employees to one another rather than to an objective performance standard, so a competent guide working alongside an exceptional colleague may appear weaker than they actually are.

Reliability problems that erode consistency

Even when an appraisal tool measures the right things, it can fail at consistency. Reliability problems show up when the same employee receives different ratings across time periods, raters, or contexts despite no actual change in performance.

Instability over time

Ratings can drift simply because the rater’s mood, workload, or external pressures change between review cycles. A manager working through a stressful peak season may rate more harshly than the same manager during a quieter month. Appraisal reliability and validity remain major problems in most appraisal systems, and new systems are often met with substantial resistance precisely because employees sense this drift even when they cannot name it.

Time-based instability is especially relevant in tourism, where business cycles fluctuate sharply. Ratings collected during a high-pressure festival season may not align with those collected during a lean travel period, even though the underlying performance is similar.

Inconsistencies among raters

In many tourism businesses, employees report to multiple supervisors – a duty manager during the day, another at night, plus a department head. When these raters use different mental yardsticks, results diverge sharply.

The technical term for this is poor inter-rater reliability. Assessment tools that rely on ratings must exhibit good inter-rater reliability, otherwise they are not valid tests. If one supervisor weighs guest feedback heavily while another prioritises operational efficiency, the same housekeeping attendant could receive very different scores in the same review window. The appraisal stops being a measure of the employee and becomes a measure of which supervisor happened to fill in the form.

Lack of training in appraisal techniques

Many supervisors are promoted into appraisal responsibilities without ever being trained on how to conduct one. They may not recognise their own biases, may misinterpret rating anchors, or may apply standards inconsistently. One study found that raters trained using a specific methodology achieved a Cohen’s Kappa value of 0.85, indicating high agreement, compared to untrained raters at just 0.5 – a striking gap that translates directly into fairer evaluations.

Training is especially important in service industries where soft skills like empathy, communication, and cultural sensitivity matter enormously but resist easy measurement. Without preparation, raters tend to fall back on instinct, which is exactly where bias lives.

Why these issues matter in service-driven workplaces

Tourism organisations live and die by the quality of guest interaction. When appraisals fail to validly capture this performance, several downstream problems follow. Talented employees feel unrecognised and disengage. Underperformers escape detection and continue to weaken the team. Promotion and increment decisions reward the wrong people, eroding trust in management. Over time, the appraisal system becomes a ritual rather than a tool – completed for compliance but ignored for decisions.

There is also a legal dimension. When ratings influence terminations, demotions, or pay, an appraisal system that cannot be defended on grounds of validity and reliability exposes the organisation to disputes and regulatory scrutiny.

Practical ways to strengthen validity and reliability

The good news is that both problems respond well to deliberate intervention. A few strategies stand out.

Standardise the criteria

Vague performance standards invite subjective interpretation. Instead of writing ‘increase sales’ as a performance standard, a clearer version is ‘increase sales by 10 percent from last year’ – a measurable target that any rater can verify. In tourism, standardised metrics could include guest satisfaction scores, occupancy contribution, complaint resolution time, or adherence to service protocols.

Use behaviourally anchored rating scales

Generic scales of 1 to 5 leave too much room for interpretation. Behaviourally anchored rating scales, or BARS, attach specific examples of observable behaviour to each rating point. BARS define scale points with specific behaviour statements that describe varying degrees of performance, making it much clearer to both rater and employee what each score actually means.

Train raters and run calibration sessions

Bringing supervisors together to rate the same sample case and then discussing why their scores differ is one of the most powerful corrective tools. Calibration exercises engage raters in discussions to ensure consistency in ratings and scores, exposing hidden differences in interpretation before they affect real employees.

Use multiple rating sources

A single rater carries a single set of biases. Multi-source feedback – sometimes called 360-degree appraisal – combines input from supervisors, peers, subordinates, and in tourism, often guests as well. Each source corrects for the blind spots of the others, producing a more balanced and stable picture.

Document performance throughout the year

Recency and primacy errors thrive when the rater relies on memory. A simple log of significant incidents, complaints handled, commendations received, and goals met turns the annual review from an act of recall into an act of summary, which is far more reliable.

What do you think? Which of these validity or reliability problems do you suspect is the most common in tourism workplaces you have observed, and what would change in an organisation if its appraisal system suddenly became genuinely fair and consistent?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://learn.saylor.org/mod/book/view.php?id=60428&chapterid=47713
  2. https://opentext.ku.edu/teams/chapter/performance-evaluation/
  3. https://www.dartmouth.edu/hr/professional_development/for_managers/performance_management/common_rater_errors.php
  4. https://txwes.pressbooks.pub/iopsychologytxwes/chapter/7-3-performance-appraisal-part-2-rating-distortions/
  5. https://factorialhr.com/blog/bias-in-performance-reviews/
  6. https://bizfluent.com/about-5445066-importance-reliability-performance-appraisals.html
  7. https://en.wikipedia.org/wiki/Inter-rater_reliability
  8. https://encord.com/blog/inter-rater-reliability/
  9. https://www.nationalforum.com/Electronic%20Journal%20Volumes/Lunenburg,%20Fred%20C.%20Performance%20Appraisal-Methods%20And%20Rating%20Errors%20IJSAID%20V14%20N1%202012.pdf
  10. https://www.numberanalytics.com/blog/ultimate-guide-to-inter-rater-reliability

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Managing Personnel in Tourism

1 Functions and Operations of a Personnel Office

  1. Characteristics and Objectives of Personnel Management
  2. Functions and Operations of Personnel Management
  3. Organisation of a Personnel Office
  4. Personnel Managerโ€™s Role
  5. Position of Personnel Department in the Organisation

2 Recruitment and Selection

  1. Essentials of Recruitment Policy
  2. The Process of Recruitment
  3. Methods of Recruitment
  4. Selection
  5. Physical Examination

3 Induction and Placement

  1. The Importance of Proper Induction
  2. Induction Process
  3. Induction Programme
  4. Placement
  5. Induction as an Integrated Part of Training

4 Staff Training and Development

  1. Defining Training and Development
  2. Training
  3. Evaluation of Training Programmes
  4. Retraining
  5. Management Development

5 Motivation and Productivity

  1. Hierarchy of Human Needs: Maslowโ€™s Theory
  2. Social Needs and Productivity
  3. Hygienes and Motivators
  4. Creating Proper Motivational Climate

6 Employee Motivation and Job Enrichment

  1. What is Motivation?
  2. Types of Motivation
  3. Theories of Motivation
  4. Motivation and Morale
  5. Job Enrichment โ€“ Meaning Nature and Objectives
  6. How to Enrich Jobs?

7 Career Planning

  1. What is Career Planning?
  2. Why Career Planning?
  3. Responsibility for Career Planning
  4. Process of Career Planning and Development
  5. Advantages of Career Planning
  6. Limitations of Career Planning
  7. What makes Career Planning a Success?

8 Performance Monitoring and Appraisal

  1. Some Activities
  2. What is Performance Appraisal?
  3. Job Performance and Performance Measurement
  4. The Problems of Validity and Reliability
  5. Methods of Appraisal
  6. Making Performance Appraisals More Effective

9 Transfer, Promotion and Reward Policies

  1. Need for a Transfer Policy
  2. Promotions and Promotion Policy
  3. Reward Policies and Processes
  4. Measurement of Performance and Reward Policies
  5. Vehicles for Rewards

10 Employee Counselling

  1. What is Counselling?
  2. Need for Counselling
  3. Counselling Functions
  4. Counsellors
  5. Skills and Techniques
  6. Types of Counselling

11 Discipline, Suspension, Retrenchment and Dismissal

  1. What is Discipline?
  2. Indiscipline
  3. Disciplinary Action
  4. Suspension
  5. Dismissal
  6. Retrenchment

12 Employee Grievance Handling

  1. What is a Grievance?
  2. Why Grievances?
  3. How to Handle Grievances
  4. The Discovery of Grievances
  5. The Processing of Grievances
  6. Steps in Grievance Handling
  7. Doโ€™s and Donโ€™ts in Grievance Handling

13 Compensation and Salary Administration

  1. Aims of Salary Administration
  2. Principles of Salary Formulation
  3. Components of Salary Administration and Pay Structure
  4. Salary Structures
  5. Salary Progression
  6. Salary Administration Procedures
  7. Other Allowances

14 Laws and Rules Governing Employee Benefits and Welfare

  1. The Concept of Fringe Benefits and Labour Welfare
  2. Objectives of Labour Welfare
  3. Statutory Welfare Provisions
  4. Voluntary Welfare Amenities
  5. Social Security: Concept and Evolution

15 Gender and Other Related Issues in Tourism

  1. Position of Women in Tourism
  2. Manager’s Responsibilities
  3. What is Sexual Harassment?
  4. Code of Conduct
  5. Conducting Enquiry by the Complaints Committee
  6. Child Labour, Human Rights, and Consumer Protection