Details in this account are generalised and figures are omitted where they would identify the client.

The situation

A business operating from three sites — a head office, a workshop and a smaller depot — connected by site-to-site links, with around forty staff between them.

The network had accumulated over roughly a decade. Switches added as the business grew, access points installed where there was a power socket, a second router installed for a project in 2019 and never removed, and cabling from at least three different installers.

The complaint was a fault at the workshop: two or three times a week, for a few minutes, everything would slow or drop. Nobody could reproduce it, the internet provider found no fault, and it had been going on for over a year.

What monitoring found

We installed continuous monitoring before proposing anything — internet quality from each site, availability of key devices, and switch port statistics.

Two weeks produced the answer, and it was not the internet connection.

Error counters on one switch port were climbing steadily. That port fed a small switch in the workshop office, which had been added by somebody at some point and served six desks and an access point.

The pattern correlated with activity in the workshop, not with time of day. The cable ran through the workshop along a route that shared trunking with mains cabling for a section, and had been fastened with cable ties tight enough to deform it.

When machinery in the workshop ran, interference and the marginal cable together pushed the link into an error state. It recovered when the machinery stopped, which is why nobody could reproduce it and why the provider found nothing.

A year of an untraceable fault, resolved by a cable and two weeks of measurement. Intermittent problems are physical almost every time, and error counters find them.

The wider assessment

Fixing that cable was a morning's work. The monitoring had also surfaced enough else to justify a proper look.

No addressing plan. Three sites using overlapping ranges, which made the site-to-site links more complicated than they needed to be and caused problems for staff connecting from home.

Everything on one flat network at each site. Cameras, the door entry system, the workshop machinery interface, guest wi-fi and the accounting workstations, all able to reach each other.

Undocumented equipment. Including the second router from 2019, still connected, still routing some traffic, and known to nobody currently employed.

No documentation at all. No diagram, no port schedule, no record of what was where.

How the rebuild was sequenced

In stages, lowest risk first, each during a planned window with testing afterwards.

Stage one: the physical faults. The workshop cable rerouted and reterminated, plus three other runs the error counters had flagged. Immediate, and it resolved the presenting complaint.

Stage two: guest separation. Isolated guest wi-fi at all three sites. No risk to anything existing, and it removed visitors from the business network.

Stage three: non-computer devices. Cameras, door entry, the machinery interface and printers onto their own segment. This took the longest, because working out what each device genuinely needed to reach required investigation rather than assumption.

Stage four: addressing. A planned scheme with distinct ranges per site and per purpose. Done over a weekend, and the most disruptive single change.

Stage five: wireless. Surveyed properly and access points repositioned, several onto new cabling. The workshop had two access points interfering with each other and a dead area at the far end.

Stage six: documentation. Diagram, addressing plan, port schedules, and credentials held where more than one person can reach them.

What we would do differently

Monitor before anything else, always. We did here, and it is the reason the project started with a correct diagnosis rather than a proposal to replace equipment that was not faulty.

Budget more time for the device segmentation. Establishing what a decade-old door entry controller needs to communicate with took considerably longer than expected, partly because its documentation no longer exists.

Do the documentation as we went, not at the end. Writing it up afterwards meant revisiting things we had already worked out.

The general lessons

Measure before proposing. A year of blaming an internet provider was a year of looking in the wrong place, and two weeks of monitoring settled it.

Check switch port error counters. They find physical faults before people report them.

Segment in stages, starting with guest access. Each stage is useful alone and none requires the next.

Our infrastructure team works on business networks across Suffolk and Norfolk — see also business network design. Start a conversation.