A telephone line delivered over IP takes one word on a quote and two very different protocols in reality. That difference explains almost every incident: the call that rings with nobody able to hear, the choppy audio at two in the afternoon, the invoice that explodes on a Monday morning. This guide separates the two flows, gives the bandwidth calculation you can redo yourself, and explains why where your telephony lives matters more than the speed of your access.
Estimate my cost →The SIP protocol, defined in RFC 3261, carries not one byte of audio. It negotiates: who is calling whom, which codec, where to send the sound. The sound itself travels in a second flow, the RTP protocol of RFC 3550. The two can take different paths, and that separation is the best diagnostic tool there is. The call does not establish at all: look at signalling. The call establishes but nobody can hear: signalling works, it is the media that is not arriving.
A call does not consume the codec's headline figure, because every packet carries its headers. With the G.711 codec, sixty-four kilobits per second of payload and one packet every twenty milliseconds, you land around eighty-seven kilobits per second per call per direction once Ethernet headers are counted. With a compressed codec such as G.729 you land around thirty-one. Multiply by SIMULTANEOUS calls, not by handsets: that is the only figure that matters, and it typically runs between a tenth and a quarter of headcount depending on the business. Thirty simultaneous G.711 calls therefore need roughly two and a half megabits per second each way, reserved, which is very little. Bandwidth is almost never the problem.
Three quantities, and none of them is throughput. First, end-to-end delay: ITU-T G.114 puts the comfortable limit at around one hundred and fifty milliseconds one way, and beyond four hundred milliseconds conversation becomes painful because speakers talk over each other. Then jitter, the irregularity of packet arrival: a buffer smooths it, but it does so by adding delay, so it spends your budget. Finally loss: beyond the order of one per cent it becomes audible. A loaded access with no prioritisation ticks all three boxes at once, which is exactly what happens when a backup starts at two in the afternoon.
Between your installation and the carrier there is nearly always a border device. It does several things at once: hiding your network topology, handling the address translation that breaks SIP, capping simultaneous calls, adapting codecs when the two sides disagree, and acting as a first barrier against fraud. Depending on its configuration it also forces the media through itself or not. Knowing what it does with your audio flow is the first question to ask when quality degrades for no visible reason.
A SIP trunk is a door onto the global telephone network, and that door is scanned continuously by robots. The scenario never changes: a weak password or an open port, a discovery on Friday evening, thousands of calls to premium destinations over the weekend, and a five-figure invoice on Monday. Three protections are worth all the rest: a spend or concurrency cap at the carrier, strict filtering of the addresses allowed to talk to your installation, and denying by default the international destinations you never call. Those three settings cost an hour and are requested from your carrier in writing.
If your installation sits in your offices, the audio crosses your internet access along with everything else sharing it. If it is hosted in a data center, two things become possible. The first is real prioritisation, because you control the link. The second is more interesting: if your telephony carrier is physically present on the site, a cross connect or a direct interconnection removes the public internet from the media path entirely. You no longer suffer a third party's congestion or a route that changes with the routing. It is the most powerful quality lever on the subject, and it costs no bandwidth, it costs a choice of site.
The comparator can tell you which carriers are present on a site, and that is precisely the information that decides your media path. It is a cross-reference of public registries and operator declarations, dated and readable as such. Trunk quality, on the other hand, appears in no registry: it depends on configuration, sizing and operations, and nobody publishes it. We measure no voice latency and we display none. What we can do is stop you choosing a site where your telephony carrier is not present.
Separate signalling from media, and half the faults diagnose themselves in a minute. Count simultaneous calls rather than handsets, and you will see bandwidth was never the subject. Watch delay, jitter and loss, because those are what people hear. Close the trunk before a robot finds it. And if you get to choose the site, choose the one where your carrier already is: it does the most for quality and it is the only decision that is hard to correct afterwards.
With an uncompressed codec and twenty millisecond packetisation, allow around eighty-seven kilobits per second per call per direction including headers, so roughly one point seven megabits each way for twenty calls. That is small: the problem is almost never throughput, it is regularity.
Because signalling and media are two separate flows. Signalling got through, media did not. Look at address translation, at filtering that passes the first flow and blocks the second, or at an announced address that is not reachable.
Rarely. Compression saves a few tens of kilobits per call and costs quality, sometimes an extra transcoding step that adds delay. On a properly sized link the uncompressed codec is almost always the better choice.
A spend cap at the carrier, a strict list of addresses allowed to talk to your installation, and denying by default the destinations you never call. Those three settings stop almost every known scenario.
The point is not compute power, it is the path. If your carrier is present on the site, the audio no longer crosses the public internet, which removes third-party congestion and route changes at a stroke. It is the most effective quality lever, and it is decided when you choose the site.
Written on 6 September 2026.
Estimate my cost → Compare data centers