Introduction

Student Evaluation of Teaching (SET) surveys:

  • can provide you with useful information to improve teaching; and
  • do provide reviewing agencies with one required form of evidence associated with teaching effectiveness as part of merit and promotion reviews (see Red Binder I-75 and APM 210-1-d).

This document provides guidance for interpreting SET reports. Please note: Understanding teaching and/or providing evidence of teaching effectiveness for reviewing agencies requires multiple perspectives/multiple forms of evidence. SETs should not be the only form of evidence on which you rely.

 

Glossary:

Student Evaluations of Teaching (SETs): Standardized surveys delivered (electronically) to students at the end of a course.

Explorance (or Explorance Blue): The platform used to deliver UCSB’s course evaluation surveys (Fall 2023-present).

ESCI: The name of the platform used to evaluate courses at UCSB from 1973-2023.

Red Binder (RB): UCSB’s official academic policy manual. Policies are delineated by section and number in the RB, e.g., RB I-75-V.

Academic Personnel Manual (APM): The University of California’s official academic policy manual. Local campus policies (like the RB) are interpretations of the APM. Policies are delineated by section and number in the APM, e.g., APM 210-.1-d.

Reviewing Agencies: Campus entities that review merit cases. For Senate faculty, reviewing agencies typically include: the department, the division/college dean’s office, the Academic Senate Committee on Academic Personnel (CAP), and the Associate Vice Chancellor for Academic Personnel. Cases for promotion and career review advancements also are reviewed by the Executive Vice Chancellor and Chancellor. Non-Senate faculty cases are typically reviewed by departmental reviewers, the division/college dean’s office, and the AVC for Academic Personnel.

SETs as Research: Investigating Your Own Teaching

Teaching is something that people learn how to do. Just as you provide feedback to your students to help them learn, SETs are one form of feedback that can help you learn.

SETs can provide data for questions about your teaching. For example, you could ask a broad question:

  • What do quantitative scores/written comments in my SET reports suggest regarding aspects of my teaching that are especially effective and/or which need attention?

You can also ask narrower questions using UCSB’s Characteristics of Teaching Effectiveness, e.g.,:

  • “What do quantitative scores/written comments suggest regarding alignment with learning goals, equity and inclusion, and implementing effective instructional strategies?”

When you are interpreting SETs to improve your teaching, consider what changes you can make in your next course(s) based on students’ feedback.

When you are interpreting SETs to document teaching effectiveness for merit review, you’ll want to use your interpretation as part of a broader account of how you’ve worked with students during this review period. This should include what has gone well; if applicable, it also should include reflections on what’s gone less well and how you have sought to improve as an instructor.


Interpreting Quantitative (Likert-Rated) Scores

In designing the required seven Likert-scale questions and one open-ended question, the campus sought to elicit feedback about how elements of teaching contribute to student learning. When interpreting these scores and/or writing about them for a merit case, this perspective can be a helpful reminder that SETs are about activities/actions rather than individuals.

2.1 Look for patterns across focus areas.

As when you interpret or analyze any data, the key is to: 1) look for patterns in the feedback; and 2) explain those patterns to reviewing agencies when you are documenting teaching effectiveness. (Remember: you need to explain how you are ‘reading’ these results!)

  • One set of focus areas is reflected in the SET questions: course design, course delivery, feedback, and environment:
    • This course advanced my understanding of the subject matter. (course design; course delivery)
    • The instructor explained course material in a way that helped me learn. (course design; course delivery)
    • This course was organized in a way that helped me learn. (course design)
    • I knew what was expected of me in this course. (course design)
    • The feedback I received in this course helped me learn. (feedback; environment)
    • I felt that the grading was fair in this course. (feedback; environment)
    • I felt that the learning environment in this course was positive. (environment)
  • It’s also possible to define alternative focus areas such as those outlined in UCSB’s Characteristics of Teaching Effectiveness:
    • Alignment
    • Equity and inclusion
    • Fostering student agency
    • Implementing effective instructional strategies
    • Attending to student growth

2.1.1 Look for patterns within your own courses and then across your courses.

Consider reading the scores in your reports in the following order:

  1. Identify any patterns within a single course.
  2. If applicable, identify patterns across courses identified as small/large/graduate.
  3. Within question categories (i.e., “Understanding, Explain, Organization, etc.).

2.1.2 Look for areas where your scores are particularly high or particularly low.

Individual scores

Mean scores in the reports provide a sense of how students experienced elements of the course represented by the questions.

When you interpret your mean scores, you can do so in three broad categories created based on an analysis of SET scores for all faculty[1]:

For any of these scores, and especially for scores of 3.5 or below, it’s valuable to look at the qualitative comments associated with the overall ratings. These can tell you more about students’ experiences in relation to the particular response. Office of Teaching and Learning instructional consultants are also here to help interpret these patterns, provide guidance, or otherwise support your teaching.

Next, look at the distribution of responses (% strongly agree, agree, etc.) and the standard deviation (e.g., +/- .8; +/- 1.4). Lower standard deviations (SD) indicate questions that were rated more consistently; higher SDs indicate questions that were rated more variably. Lower or higher SDs can tell you whether your specific scores are within the range of the averages for the majority of scores. Analysis suggests that if the difference between the instructor’s mean score on a question and the normative score for the question is 0.5 or more, it’s worth investigating.

Finally, focus again on patterns. For example:

  • Are the ratings, overall, within the same range (i.e., are most within the range of a 4 or 5 [agree/strongly agree]? Are most ratings in the range of a 1 or a 2 [strongly disagree/disagree])?
  • Are there outliers in the pattern, i.e., one or two questions that are rated lower or higher than other questions?

It’s important to attend to written comments even when scores are high, but especially important for questions that have consistently low scores, outliers, and/or when the difference between instructor and normative ratings are 0.5 or greater. Comments may provide more information about those differences so that instructors can reflect on their teaching, and so that they can provide more context for reviewing agencies during merit review. Figure 1 shows Professor Snowy Plover’s scores for three small courses. (Note that Professor Plover’s report is a facsimile.)

 

Reading the Results: Professor Snowy Plover’s Initial Interpretation

Figure 1. Professor Snowy Plover’s Report

Professor Snowy Plover just received their first summary report and is preparing to write a teaching statement for their merit case. They begin by reviewing individual course ratings. They note that the scores are relatively consistent in each course, except for a 3.9 rating on “Explain” (i.e., the question that asked students to indicate agreement with the statement, “The instructor explained course material in a way that helped me learn”) in SPR 33 and a rating of 3.2 on “Understand” (i.e., “This course advanced my understanding of the subject matter”) in SPR 125. Professor Plover decides to investigate this by paying especially close attention to students’ qualitative feedback for these courses to see if comments provide additional information about what students’ appreciated about the other aspects of the courses and what they noted might have gone differently about feedback.

Next, Prof. Plover reads across all the courses in the report. They see that the scores for SPR7 seem consistently higher than other courses. Prof. Plover implemented some new strategies to help students learn in the course. They make a note, thinking that they might want to highlight some of this work in their teaching statement.

Finally, Prof. Plover looks down each column to identify any pattern in the score report. Here, too, they see an increase in student ratings from Fall ‘23 to Spring ‘24. They note this, as well, as possible evidence supporting an analysis of more focused efforts to improve student learning via teaching.

 

2.2 Small courses may require additional evidence.

Statistics are meaningful only when the number of responses is  large enough to render them meaningful. We’ve already done a statistical analysis of SET scores to create the comparison norms for departments and the campus -- this analysis is behind  why a “small” class is a class with 40 or fewer students and a “large” class is a class with 41 or more students.[2] But if you’re teaching a very small class -- anything under about 15 or so -- statistics may be less meaningful. For very small classes, it will be even more important to have other forms of evidence to support your evaluation and discussion of teaching effectiveness. These other forms could include syllabi, your teaching statement, a course observation, annotated assignments, or other forms of evidence described in the UCSB Teaching Effectiveness Guidance (p.2).

2.3 Consider response rates

Response rates to online SETs can be a challenge at UCSB -- in fact, they can be a challenge at many institutions of higher education right now. Experience at UCSB has shown that students are more likely to complete SETs when:

  • Instructors talk with students about why SET responses help them improve their teaching. (“I hope you’ll complete these -- your feedback is really important to me, and I review it carefully. Students comments have really helped me shape my teaching….”)
  • Instructors give time in class for students to complete their evaluations. (“I know your lives are crazy busy, so I’ve built some time in so you can complete your evaluations. I’ll step out…”)

When you’re analyzing your SET scores, be sure to attend to the response rates, which are included in evaluation reports. Low response rates make generalizing feedback difficult. Research shows that:

  • Classes <40: require 95% response rate for a 95% confidence level; 40% for an 80% confidence level.
  • Classes >100, 87% response rate for a 95% confidence level; 21% for an 80% confidence level (Nutly, 2008).

In this study, a lower response rate for a medium to large course (in the 41-99 range) was found to be analogous to higher response rates. However, in small courses, low response rates can greatly influence the mean and may not correctly characterize the data from all students. Put differently,  if a small percentage of enrolled students submit surveys, it may lead to a question about response bias. You could ask: “what might be different in the responses of those students who chose to complete SETs and those who did not?”  Low response rates are another very good reason to make sure you have multiple forms of evidence to gather information about and provide evidence of teaching effectiveness.

2.4 “Contextualize” your SET responses.

Reviewing agencies repeatedly emphasize the need for faculty to “contextualize” their analyses of teaching (effectiveness). Contextualizing SET results is also important as you interpret responses for your own improvement, as well.

“Contextualizing” involves providing information about the circumstances in which the course is taught. While these circumstances are many, especially for merit reviews you will need to provide an analysis that is concise, and that readers who are not in your discipline can understand.

Some contextual factors you might want to consider are listed below. Note that this list is extensive, and not all of these elements will apply -- you’ll want to focus on the ones that seem most important, and be sure to explain how and why they are important.

Here, too, think of categories. Some factors are institutional, i.e., the place of a course in the curriculum; the structure of the course; the setting (physical or virtual) in which the course is taught. Others are person-related: the identities and experiences that students bring to the course; your own identities and experiences. There are more and you should of course add your own, keeping in mind that focus and explanation for reviewing agencies are key.

2.4.1 Institutional Contextual Elements

GE, Lower Division, Upper Division, Graduate Courses

Students might bring different kinds of knowledge, investments, and/or motivations to a course depending on where it falls in their learning trajectories. As instructors learn about these contextual factors, they also learn to adapt their instruction.

Is this…

  • a general education course?
  • a lower division or upper division course?
  • an elective or required for the major?
  • an introductory, advanced, or elective graduate course?

Course structure.

Designing a course where people can learn requires taking the course structure into account. A large course with sections is one structure, a large course without sections another, a large course with reader(s) yet another. As instructors experience different course structures, they learn to design for learning.

Is this…

  • a large course with sections and TAs, or a large course with no TAs (but readers)?
  • a small seminar?
  • a workshop, production, or performance course?

Setting for the course.

Designing for learning involves considerations of where the learning is taking place. A lecture hall with “turn to team” seating is quite different from one where seats are bolted to the floor; a project-based learning room is different from a room where there is little space between desks. A synchronous online class with structured breakout rooms is different from one that is “only lecture.” As instructors learn about teaching spaces, they consider how these spaces enable different types of learning. Some factors that might be worth considering, for instance, are whether this course is:

  • offered in a room or environment that is conducive to learning?
  • a hybrid or online course?
  • The first time you’ve taught the course? The ninth time?
  • Early, middle, or late in your career as a faculty member?
  • A GE course, a course in your major, a graduate course?

Each of these factors (and likely others) might affect the expectations that students bring to your course or their experience of it. When you analyzing your SET responses, it’s helpful to take all of these into account. When you are interpreting your analyses for reviewing agencies, it’s essential to lay out relevant contextual elements for them. Remember: while these elements may be familiar to you and part of the story of your analysis, you need to lay out these out for people not familiar with you or your courses.

2.4.2 Comparing to department, college, and campus norms.

When comparing individual scores to department, college, and campus norms, it’s helpful to take the context of the course into consideration (see 2.5). Large courses that fulfill GE requirements, for instance, are best compared to college norms, because this number includes the totality of scores across all majors, just as GE courses include students across majors.

Case Study: Professor Plover’s Course Context

Interpreting the SET scores, Professor  Plover grouped the ratings in two chunks: course delivery and course design, feedback, and environment. Plover saw that overall, students generally agreed that their delivery helped them learn. This made sense, Professor Plover thought -- they had heard repeatedly that they were a very engaging lecturer, and this was borne out in students’ written feedback, where students said Professor Plover was entertaining, fun, and lively. But Professor Plover observed that students said the course design, feedback, and environment were less helpful for their learning.

Professor Plover also knew that the room where the course had been scheduled had made group work especially hard, too -- the chairs were bolted to the floor and it was hard for students to talk to each other. They decided that next time, they would request a large project-based learning room so students could collaborate more easily.

Professor Plover saved these notes to a “teaching statement” document that they could use for their upcoming merit review. When they wrote their teaching statement, they would be able to provide context: this was their first time teaching a large class; some things went well, and they knew what they would work on with their teaching mentor and Office of Teaching and Learning to develop their teaching as they taught the course in subsequent offerings.


Interpreting Student Comments (Qualitative Data)

All instructor SET surveys include one open-ended question, “Please provide any additional feedback for the course instructor about your learning.” They also include boxes for brief comments adjacent  to each Likert-rated question (e.g., “This course advanced my understanding of the subject matter”). Ideally, students’ comments will focus on teaching. Guidance is provided here for when that is the case, as well as for when it (unfortunately) is not.

Before we provide this guidance,  a word on language. Just as our language is powerful for students, we want to acknowledge that students’ language can be powerful for us. Many instructors have received comments from students that have cut close. Especially when these comments are anonymous, comments can lead us to wonder: “Who? Why? What did I do to provoke this?” As difficult as it might be, though, we hope that you’ll take a half-step back from comments on SETs, thinking about them not as a commentary on you, but as data about students’ experiences of their interaction with the entirety of the course environment. This perspective might be helpful as you analyze the comments; it also might be helpful when you draft your teaching statement for merit review.

3.1 Teaching-Focused Comments

3.1.1 Read Globally

First, read through all of the comments. Here, you might imagine that you’re getting a sense of the terrain: what’s there, overall?

3.1.2 Look for patterns

After your global read, start to identify any patterns, keeping in mind a question as you go: What do groups of student comments tell you about students’ (plural) experiences of their interaction with the entirety of the course environment? If you have comments associated with specific Likert-rated questions, group your reading in the same patterns in which you grouped your interpretation of the quantitative ratings (see 2.1, 2.2, 2.3).  After you’ve analyzed the comments on Likert-rated questions, turn to the comments on the open-ended questions. (e.g. “Please provide any additional feedback for the course instructor about your learning.”)

 

Case Study: Professor Plover Reads Their Comments

Analyzing their Likert-rated (quantitative feedback), Professor Plover saw that students found their delivery helpful for learning. Prof. Plover also recognized that lower student ratings in the areas of course design, feedback, and environment signaled areas for more attention.

“It’s the comments that will tell me more,” Plover said to themselves. First, they skimmed all of the feedback. This was a large course, so Plover had a few comments from boxes next to the Likert-rated questions and a large number of responses to the last question.

Almost immediately, Professor Plover saw some anomalous responses. “BEST COURSE EVER!” was one. Another read, “Whaaaat was this course? I didn’t get a single thing! The whole thing was a mystery.” Professor Plover recognized that neither comment provided much detail that they (or reviewing agencies) could use to understand the course, so they weren’t as helpful. They also were atypical.

In this instance, Prof. Plover decided, then, to put those comments aside. Instead, Prof. Plover focused on recurring comments about course delivery and course design, feedback, and environment. They were careful not to cherry pick, realizing that a forthright and honest analysis would help them learn more about what was going on. They knew, too, that this kind of appraisal would indicate to reviewing agencies that they were paying attention to their teaching -- something recognized in the characteristics of teaching effectiveness and valued by reviewing agencies.

Paying attention to students’ language, patterns began to emerge: Students said they didn’t feel they had enough information about what the big ideas were in their discipline, or how they were supposed to learn those ideas, and that it seemed like Professor Plover thought they knew these ideas already. They didn’t feel like they were getting the feedback they needed to ask about those ideas, either -- and they weren’t comfortable asking questions in class.

Professor Plover sat down with a thud. They couldn’t believe it. They had tried so hard in this course -- they had poured a huge amount of energy into it and this was the result? They wondered: “what were these students thinking? How could they not get it?” But the moment that thought entered their head, Prof. Plover quickly stopped. “WAIT!”, they thought. “This is exactly the feedback I need to help me understand what students need to help them learn.” After a moment’s reflection, they also realized: this was their first time teaching this course. Their department chair and teaching mentor had both reminded them that any first time was just that, and that they would learn from the experience. Reflecting on that learning, identifying what they might want to keep and do differently next time, would signal that growth, probably make the class better, and be important to point out in their teaching statement for merit review.

After their moment of introspective revelation, Professor Plover scheduled a meeting with an Office of Teaching and Learning instructional consultant to puzzle through this experience. After a cheerful and helpful chat, they scribbled some questions to ask as they put together their next offering of the course:

  • “How can I provide more explicit expectations for learning in this course/discipline?”
  • “How can I provide more information about what a “good job” representing learning [via writing, tests, or other assessments] looks like for students?”
  • “How can I include ways for students to work together to practice with what they need to do to succeed and provide formative feedback?”

 

3.1.2 Low response rates and comments

In 2.4, we explained that the quantitative data from very small courses may not be as useful the data from large courses with a larger number of responses. Although qualitative comments are not statistically-driven, they still may provide some helpful feedback. Here, too, the key is to look for patterns. If the class is small and each student’s comments describe a different experience, you might ask: “Could I reinforce a more consistent theme/pattern in the course?” If a few students experienced something more or less strongly than others, you could ask, “What was my intention? How did some students access that, and others not access that?”

3.2 Comments Not Focused on Teaching

The category of “non-teaching focused comments” includes (but is not limited to) comments associated with one’s identity or person rather than a student’s experience of the course (e.g., comments on appearance, multiple identities, tone of voice, religious expression….) These comments can be tied to  Likert-rankings or be distinct from them. For instance, a student may indicate that their rating for the question, “The instructor explained course material in a way that helped me learn” was because they didn’t like the instructor’s tone of voice. Alternatively, while the final open-ended question (“Please provide any additional feedback for the course instructor about your learning”) has been framed to focus on “your learning,” students may write anything they would like to in this space.

Before we provide guidance on non-teaching focused comments, we want to acknowledge that these can be hurtful, frustrating, both, or many other things. They may be hurtful because they compound harm that has been done by countless others’ words and deeds; they may be frustrating because, as part of the merit review process, you might wonder how others will read these comments. Additionally, because comments are anonymous and directed at instructors, there’s no way to engage in a dialogue with the people who made the comments about their words.

3.2.1 You ’re not alone: you have resources for support.

When you receive non-teaching focused comments, you don’t need to figure out what to do with or about them on your own. Talk to your department or division teaching mentor, your department chair or vice chair, and/or an Office of Teaching and Learning Instructional Consultant. We are all here to help.

3.2.2 Framing harmful comments: Responsibility and UCSB’s Principles of Community

Sometimes, instructors receive comments related to identities, characteristics, features, or other aspects of one’s person -- rather than teaching. Harmful language is used by people and beliefs are held by people -- but as members of a community, the campus has a responsibility to engage in dialogue about defining, maintaining, and participating in collaborative beliefs. Currently, these are outlined in UCSB’s Principles of Community. If and when you choose to address harmful comments included in your SETs, it might be helpful to frame these comments within the Principles, noting where they do not (or, potentially, do) align with the Principles. Just as the campus seeks to cultivate actions guided by the Principles of Community and does not tolerate actions that fall outside of them, so comments that are not aligned with the principles will not be taken into consideration in the merit review process.


SETs and Epistemological Assumptions

Finally, a note on SETs and epistemologies. SETs are controversial. There is research showing that the scores and comments that are submitted via these anonymous commenting opportunities can reflect all kinds of bias -- sexism, racism, raciolinguicism, and many others. This is why the APM and RB require at least two forms of evidence to demonstrate teaching effectiveness as part of the merit review process. From 2021-2023, a group of faculty at UCSB, the Teaching Evaluation Workgroup (TEW), conducted extensive research to revise the course evaluation questions used for evaluating teaching, designing questions that focused on an instructor’s teaching, rather than the instructor themselves. The TEW’s work was preceded by the work of three other committees (two systemwide, one campus) investigating issues associated with teaching evaluations. As a system and as a campus, UC/UCSB are aware of these issues.

Validity/Efficiency

At the same time, there is always a need to gather student feedback on teaching -- both for instructors to improve their teaching, and for reviewing agencies to understand teaching effectiveness. Complicating this desire are the many involved in this process. Faculty teaching large courses, for instance, need to gather data from hundreds of students -- sometimes many hundreds during a single term. Faculty teaching smaller courses may be able to gather deep data from a few students. And reviewing agencies, who analyze hundreds of cases each year, need to be able to understand multiple interpretations of these data and conduct their own analyses. Hence, there is always a tension between validity (gathering/including data that represents the fullness of one’s teaching) and efficiency (analyzing data in valid ways as quickly as possible) built into the evaluation process.

Quantitative/Qualitative

Quantitative scores on SETs are underscored by the presumption that it is possible to provide a stable, “objective” ranking of characteristics associated with teaching. This presumption itself is rooted in an epistemology that presumes that there is an objective reality that can be obtained, and by “controlling” for different characteristics or biases associated with rankings, that objective reality can be represented.

Qualitative comments are underscored by the presumption that what is measured has a bearing on which sort of evidence is preferred or more informative. They’re also underscored by the principle that people experience and create meanings that begin to constitute the realities in which they exist. These are sometimes described in words. In this perspective, there is a presumption that if we gather multiple perspectives on an experience and look for patterns, we can begin to learn about whether there is consistency among understandings of that experience.

The different presumptions and perspectives that instructors bring to quantitative and qualitative data are significant for interpreting SET scores. Some faculty believe that quantitative scores are “meaningless”; others believe that they represent reality. Similarly, some instructors believe that only student comments matter, while others believe that these comments are not meaningful. Many instructors, too, are somewhere in between these perspectives (or hold others). As you consider your SET scores, it’s helpful to approach these as one -- and only one -- form of evidence that can help you improve your teaching and/or document teaching effectiveness. The “second form” required by RB and APM is yours to choose. If you find that SETs do not reflect your epistemology, select a form of teaching evaluation that does -- because you have agency to do so.

Office of Teaching and Learning consultants are available to help faculty interpret and act on SET reports or provide additional information about interpreting these documents. Please contact OTL at otl-info@ucsb.edu if you’d like more information or would like to chat with an instructional consultant.