7 years of self rebooting radios

I just noticed one of my epmp3000 self-rebooted, again. I have been using these APs for 7 years and this problem has existed the entire time. I have never found a firmware version that does not do this. The only APs that do not self-reboot are ones with small number of clients. As soon as the numbers get over around 20 or so, it begins. The APs can run for up to 100 days, but sooner or later they will self-reboot. My installations are perfect and high quality cabling used, with grounding etc - it is not my installs at fault. I don’t usually get to see the cause, as the reboot clears the log, but occasionally I have caught it, and what I see is most or all stations have dropped, then the radio self-reboots.
Its always the same - the radio is fine for a year or so, but as soon as station numbers exceed 20 or so, the self reboots begin. They don’t happen regularly, but eventually they always happen. I run 80Mhz channel width as there’s no noise, and keep firmware updated. 5.12 just did it after running for just 1 week. Its really annoying as it obviously dumps all customer traffic, and we are all trying to keep Elon out - this stuff does not help. Please do better for us Cambium.

I am running 8 3000 all running 25-30 clients, I have seen an occasional client drop maybe 4-5, but not all, I don’t update my firmware until bugs are worked so I have see a year+ on a few ap’s I usually reboot most of them around 200 days , do u have more than one ap doing this or just the one? If you know it is going to happen at 100 days, maybe just do a manual reboot at 50-60 days at about 4:30 in the morning, that’s when my traffic is the lowest, and I do maintenance, when needed

I have ~60 ePMP APs (3000 and 3000Ls). And yes, they indeed do that. And to be honest - quite oftenly. I have a monitoring agent that messages me when client count on AP sharply drops to some low number (this could indicate AP crash or city-wide power outage of course) and when AP itself becomes unreachable, so I count quite many of these AP crashes. I have collected lots of useful data and probably know why this is happening, communicated with cambium for few years about this, they are aware of this problem, but haven’t done anything about it and they also lost further interest to investigate, tickets got silent. I lost hope and haven’t invested a cent in 4000 series yet :slight_smile:

31 AP’s and have provided info here and directly to cambium. between 160 and 220 days you will have a reboot. I have only had one last more then 220 it made 228. I have equipment at the same site with over 900 days uptime. All of our sites have minimum of 8 hours backup power.
I will say compared to my similar number of LTE sectors the e3000 is rock solid. but compared to my PMP450 was over a year prior to a firmware update and my Tarrana is at 541 days as of now. So there is a bit of you get what you pay for other then the LTE is really expensive garbage.

I can’t say I don’t agree with you. That is why I stopped investigating and digging deeper. I just accepted that it must have been part of price for our only available path from our former MT based network. It’s just hard for us when this is the least stable part left in our whole network, the remaining being trouble-free GPON.. And of course indifference of Cambium themselves. They could at least tell us “sorry guys, we know this problem, but we’re not going/can’t fix it”.

As for the problem itself, it works like this: it’s totally random for us, there is no clear pattern for this to happen. Some APs crash after hundreds of days of uptime, some can crash three times in one week. It has something to do with RF environment, because there are ones there that have never crashed for me. Most likely reason is that AP receives some malformed air frames which then eventually crashes the wifi driver, all clients drop (I sometimes react in time to log-in into the AP and reboot manually to save downtime) and software watchdog recovers by complete reboot of AP. So nothing really to be done by you here. And because of this random nature, I believe it’s not that easy for Cambium to collect data neccessary to fix it. Or they know the exact problem, but can’t do anything because it’s not in their power - if they’re using some ready-made buggy atheros/qualcomm drivers, or if there’s some hardware limitation in the chipset.

What environments are you operating these AP’s in? High humidity? Prone to lightning storms? Near the ocean?

We operate in a mostly dry (high desert) environment and have not seen any of the issues you folks are describing. We’ve been deploying ePMP since the original e1k came out and every version since then. We have quite a few AP’s with more then 20 SM’s… we have a few with around 50 working fine. We typically run the newest firmware available. ATM, we’re running 5.12 on all of our e4k radios, and 5.11 on all of our legacy ePMP.

I’m not trying to discount your experiences, just trying to see if there’s a difference between our deployment conditions and yours.

Climate has absolutely nothing to do with it. We operate in standard moderate European climate, and this happens in summer as much as in winter. Client count per AP is also not a meaningful factor - APs with 11 clients crash as often as ones with 55. One thing I suspect might be contributing to this is GPS pulse stability. We have had lots of GPS troubles because of infamous second harmonic issue from nearby LTE base stations. Lost sync might worsen the calculations chipset or software has to process. But this doesn’t explain why this also happened on sites where there are no LTE stations anywhere near and sites themselves operate in area isolated low in valley surrounded by forest foliage. 5.12 doesn’t help at all. I don’t think much is being done in 5.x for 3000-series. What in fact was done - at approx 5.10.x they did something with this GPS sync loss issue, where GPS doesn’t self recover until you reboot the AP, though it fixed one problem, but introduced others..

Do you folks use PPPoE? Anyone use LLDP? Do all radios have some sort of firewall protecting them from the broader internet?

Just a few notes on our ePMP network…

  • We don’t use PPPoE
  • We don’t use LLDP
  • AP’s use private IP addresses
  • All SM’s with public IP’s and using NAT have common mgmt ports firewalled upstream
  • All radios have every service and option disabled with the bare minimum needed to operate
  • We do not use traffic shaping on the radios (but we do have the QoS settings enabled)
  • IPv6 is disabled
  • AP’s are set to bridge mode and the vast majority of SM’s are set to NAT mode (there are a few set to bridge mode in order to provide static IP’s to select customers)
  • We use fixed mode TDD with a 75/25 duty cycle 5ms and GPS sync for all AP’s. Most AP’s have the option of both onboard GPS and ethernet sync injection

Here are some example screen shots of how most radios are configured:

I have lodged tickets oevr the years but they don’t go anywhere. I get asked to send support files and do, but then don’t get a reply. I think its in the too hard basket and no one can be bothered. To be fair, its hard to get support files that show anything from a radio that has just self-rebooted.

My config is similar to Eric’s, accept I run 80Mhz channels. I thought that would be it, but I doubt the others with the same issue are also running 80’s.

It sux when you spend big money on backhauls to make your network rock solid, because your F400c links kept dropping, only to be let down by the actual APs themselves.

Mine self-reboot after 1 week or 3 months. There’s no LTE or noise at all where most of them are. GPS seems fine. I usually get around 60 days, but posted this in anger after updating to the latest fw and getting a reboot after just one week - aaarrgggg!!

I have a few 3000L radios and they seem to run for ages without self-rebooting, but swap one out for a 3000 and its guaranteed to crash evetualy. The only constant I can see if radios with under 10 clients run much longer, and the 3000 crash much more than the 3000L, no matter the client numbers.

My wisp covers beachy areas and hilly areas. The ones on the beach do crash more than others, but they have the highest number of clients. Also, there’s a 3000L at the beach that goes ages without crashing, and a 3000 close by with low client numbers that also doesn’t crash as much.

It would be good if Cambium would get serious about helping us with this. If there’s a few that saw this post suffering from it, there will be plenty more that didn’t see this post.

riddle, if there’s no LTE, your GPS is fine, and a 3000 crashes while the 3000L at the same site stays up - that probably rules out climate, channel width (20/40/80 - doesn’t matter, we only run 20/40 tho), client count, GPS, and your installs. The one variable left standing is the AP itself: 3000 vs 3000L. I’ve got ~60 of both and I’ll concur - the 3000L doesn’t crash the same way. It has its own issues, but nothing this severe.

So having your input I’d like to walk back my GPS theory a bit. A closer look at my own crash logs (you said they’re useless from a self-rebooted radio - they’re not, in fact they hold a lot) convinced me GPS isn’t the main reason. In my crashlogs the reboot usually happens while GPS is fully locked and healthy - long steady streaks of good pps right up to the crash, not a flapping mess. I thought my bad GPS environment might make it more frequent, but it’s clearly not the only cause, because people with clean GPS get hit too, obviously. However our common factor is ePMP3000 AP with running GPS/Cambium Sync. And whatever would happen to GPS sync, shouldn’t send the AP swimming.

Eric, I think you’re not experiencing it mostly because of a mix of luck and a calmer RF environment — not because your config is somewhat “correct”. You also probably do not have that many 3000 APs left? Swap a busy 3000 into one of the noisier sites and I’d bet it starts at some point.

On the cause itself - the thing that trips the radio is different almost every time. Sometimes the WiFi chip just silently hangs and stops clearing its own internal queue, sometimes it hands the host a malformed frame it should`ve dropped. I could get into detail on any of it, but whatever, the point is there’s no single bug here, there are many. And every one of them ends in the exact same place: instead of just resetting the one radio module, the software makes a deliberate decision to panic and reboot the whole AP.. Meanwhile, what’s interesting, elsewhere in those same logs the firmware hits the same kind of stall and quietly recovers - so the AP clearly knows how to survive this, it just sometimes picks the most unpopular option instead. So “it’s totally random/no pattern” is only half true. The trigger is random all over the place, the crash is not - it’s the same overreaction to a hiccup. Which means the fix was never about chasing the one responsible bug, it’s about changing that one reaction - and that’s squarely in Cambium’s hands to do.

That’s what makes the silence hard to swallow. This isn’t a mystery that needs a hundred more TSFs to solve. I’m not sure, but Qualcomm’s newer WiFi code probably already handles this “queue full” situation by recovering instead of killing the AP. Cambium’s 3000 branch just never picked it up, and on top of that they made their own call to treat the hiccup as fatal. Maybe there’s a legitimate reason for that - if so, I’d genuinely like to hear what I’m trading then. But when they ask us for yet another support file, it kind of misses the point: the reaction is the problem, and they can already see it in any one of the files we’ve all sent them.

As for why they do this.. The 4000 uses a completely new chipset on a QCA codebase where this problem is most likely ironed out, so the fix path for the 3000 exists conceptually, they’ve just chosen to move forward rather than back-fit it. If more of us show the same crash across totally different countries/climates/configs, their “send us a support file” stops being a reasonable answer.

TBH, as I’ve said, this is now affecting our buying decisions. The crash itself is what we could almost live with, it’s the years of silence after sending the TSFs to nowhere that has pushed us to pause 4000 plans and scale back our wireless expansion. I can’t keep building on something that might get taken down by some unique mystery they won’t acknowledge or fix. We’d rather not, but it is what it is. A simple “yes, we know, and here’s whether it’ll be addressed” would go a long way.