What Is MOS (Mean Opinion Score)?

Also called: MOS score

Related problems: Users complain about call quality but we can't measure it; Choppy or robotic voice on some calls and not others; Provider says the network is fine but calls still sound bad; No way to prove call quality to a vendor or in an SLA review

Mean Opinion Score (MOS) is a rating of perceived call quality on a scale from 1 (bad) to 5 (excellent). It started as the average score that panels of human listeners gave to test recordings, and that is still how it is defined. Today most MOS figures that buyers see are estimates: phone systems, session border controllers and monitoring tools calculate them from measurements of packet loss, jitter, latency and the codec in use, to predict how a call would sound to a person.

At a glance

  • A 1-to-5 scale for how a call sounds to listeners; higher is better.
  • Most reported MOS figures are estimates calculated from network measurements and codec, not live listening tests.
  • As a rough guide, around 4 or above is good and below about 3.5 tends to bring complaints, but the ceiling depends on the codec.
  • Useful for spotting and comparing call quality problems; only meaningful if you know where and how it was measured.

What problem it solves

Call quality is subjective. A user says a call “sounded bad”, the IT team checks the network, and the provider says everything is fine. Without a common measure, the conversation goes nowhere.

MOS turns perceived quality into a number that can be tracked, compared between sites, carriers and time periods, and used to trigger alerts. Because the estimated score is driven by measurable network conditions, it also points toward the cause: a drop in MOS that lines up with rising packet loss on one site’s link tells you where to look. For buyers, MOS is a way to check whether a voice provider, SIP trunk or network change is actually delivering the quality users experience.

How it works

Listening tests. The original method, standardized by the ITU, plays recorded speech samples to listeners who rate them from 1 to 5. The average rating is the MOS. It’s accurate but slow and expensive, so it’s used mainly for testing codecs and equipment.

Estimated MOS from network data. Most VoIP systems calculate MOS from call statistics. A common approach is the ITU-T E-model, which combines the codec, delay, jitter and packet loss into a rating factor (R-factor) that maps onto the MOS scale. These figures appear in phone system dashboards, call detail views and network monitoring tools.

Audio comparison. Some test tools place calls, record what arrives, and compare it with the original audio using standardized algorithms. This captures audio effects that network statistics miss, but it usually runs on test calls rather than on every real call.

Codec ceilings. Each codec has a maximum achievable score. Narrowband codecs used on traditional phone calls top out lower than wideband (HD voice) codecs, so a “4.1” on one system may not mean the same as a “4.1” on another.

Where it’s measured. A MOS taken at the provider’s edge reflects their network; one taken at the endpoint reflects the user’s Wi-Fi, headset and local network too. Both are useful; they answer different questions.

Call quality reporting is a common differentiator between providers; see our Unified Communications as a Service page for how to compare them.

When it matters for buyers

  • Moving to a cloud phone system or SIP trunks. Baseline MOS before and after cutover so you can prove whether quality changed.
  • Persistent complaints about call quality. MOS trends by site, carrier or device help find the cause.
  • Selecting or renewing a voice provider. Ask what quality data they expose and whether any of it is in the SLA.
  • Supporting remote and hybrid workers. Home networks and Wi-Fi drive many quality problems that only endpoint-based scores capture.

Questions to ask vendors

  • Do you report MOS per call, and is it estimated from network data or measured from audio?
  • Where is MOS measured: at the endpoint, at your edge, or both?
  • Which codecs do you use, and what is the maximum MOS each can achieve?
  • Is there any call quality commitment in the SLA, and is it expressed as MOS, packet loss, jitter or latency?
  • Can we export call quality data or receive alerts when scores drop?
  • How do you help troubleshoot low scores that trace back to our network or the user’s home connection?

How it differs from QoE

Quality of experience (QoE) is the broad idea of how good a service feels to the people using it, across voice, video, applications and more. MOS is one specific, standardized metric for voice (and, in related forms, video) quality on a 1-to-5 scale. In practice MOS is one of the main ways to put a number on voice QoE, while quality of service (QoS) is the set of network techniques, such as traffic prioritization, used to keep MOS and QoE high.

Frequently Asked Questions

What is a good MOS score?
As a rough guide, scores around 4 or above are generally considered good, scores in the mid-3s fair, and scores below about 3.5 tend to bring complaints. The ceiling depends on the codec: narrowband codecs top out lower than wideband ones, so compare scores measured the same way.
How is MOS measured on our calls?
Phone systems, session border controllers and monitoring tools usually estimate MOS from measurements of packet loss, jitter, delay and the codec in use, using models such as the ITU-T E-model. Some tools compare a recorded test call with the original audio instead. Few organizations run true listening tests.
Can MOS be higher than 5?
No. The scale runs from 1 (bad) to 5 (excellent). In practice estimated scores rarely reach 5, because codecs and networks remove some quality.
Is MOS included in VoIP SLAs?
Sometimes. Some providers commit to a MOS target on their own network or between their equipment, but many SLAs cover packet loss, jitter and latency instead. Check where any MOS figure is measured and whether it includes your local network and the internet.
Why are our MOS scores good but users still complain?
The score may be measured at a point that misses the problem, such as the provider's edge rather than the user's headset or Wi-Fi, or it may average out short bursts of bad audio. Echo, volume and background noise also affect experience without always showing in a network-based estimate.

You Don’t Need Another Sales Call. You Need an Answer.

30 minutes. No pitch. Just an honest conversation about where you are, what you need, and whether working together makes sense.

We use your details to set up and prepare for the call, and send the newsletter only if you ask for it. Privacy policy.