SHimmer45 Posted June 13, 2014 Posted June 13, 2014 Quote Originally Posted by Stuclark View Post I suspect it's a nic driver / firmware issue (as others have said) - have you got any other nics you can put in the server and use instead of the current one? No, I haven't. The NIC drivers have been updated and the server does seem to run better but that hasn't changed the issue, sadly if it appears to run better with newer drivers it could be a NIC that is going wonky and the newer driver is coping with the flakeyness better for the sake of trouble shooting i would be purchasing another NIC (Intel would be a good bet) and running with that if your server hasnt got a second NIC. If you have any managed switches you can look at the data counters for each port to see if you are getting excessive traffic over one port or several, and i wouldnt rule out an iffy switch possibly. Are the clients able to ping the server when they are not able to right/read from the network share? thinking a little bit outside the server you havent got some child who is looping a network point somewhere in the building which isnt close the server so it takes a little while for the loopback to be felt across the network.
witch Posted June 13, 2014 Author Posted June 13, 2014 No managed switches, sadly. I dont think I said that when it happened the children trying to log on couldnt - it got stuck on "user profile" I think. I've no idea if clients can ping the server when the "freeze" happens. Problem is, I am not here often when it happens. I will try and sort this out. No loopback - I've checked every one!
Bananas Posted June 13, 2014 Posted June 13, 2014 That's the trouble with short weekly etc visits, been there done that. It's going to be mighty tough to diagnose without being able to do some tests with it "frozen". Verrrry tricky one, good luck with it
Michael Posted June 13, 2014 Posted June 13, 2014 The one thing that has changed recently is temperatures in the UK. We're now pretty much in summer season and servers without proper ventilation can offer random symptoms. Ideally the server should be connected to a UPS, however in these circumstances, try completely shutting down the server and removing the power lead for 5 minutes. Whilst it's off, give it an inspection internally and check it's clean. Now switch it back on and see how it goes. From what you're saying it doesn't sound like a switch issue. Normally a dying switch will either bomb out completely and display red (in the context of HP's), or you'll find that ports randomly start dying for no reason.
TheRobins Posted June 13, 2014 Posted June 13, 2014 Ok been reading this thread with interest, what sticks in my mind is the fact you can tie it down to a Monday at 1.20, that doesn't say traffic/hardware failure to me. Now I am going back a long long time now, I used to have the netgear DG834 home modem/router. There was a large bug within it where by once a week at a random time it would request a NTP time update from the internet. At the time I was using OpenDNS which for some reason block access to the netgear NTP server, once a week at the same time our internet would drop out and re-connect. Of course this isn't the issue here but I would be interesting to see the following info Switch make and model In event viewer do you have any tasks scheduled for that time, some SIMS plugins that request data look in at that time Do you have any "odd" devices on your network, I have seen all sorts from boiler management systems, door entry systems and the like. Have you also get a spare pc you could use, Place it on the network. Join the domain, create a share, and then map it to a pc in the school. When your shares from the main servers die, see if your spare pc's share is still open. If it is, you could basically pin point it to the server.
witch Posted June 13, 2014 Author Posted June 13, 2014 The server has a UPS I have done the switch it off for 5 mins bit and when I replaced the USB cable to the backup I did an internal inspection and it was clean. As I said, doesnt always happen at 1.20, in fact I think it may happen earlier as it did on Tuesday and it is only when the class logs on it becomes apparent. I knew earlier on Tuesday as a teacher happened to be doing something. The switches, of which they are 4, are DLinks but I can't say the model offhand. No weird events in event viewer and no SIMS or other MIS on my server (we have two networks). No odd devices either - they are all on the admin network. I like the idea of the share creation - will have a go at that - thanks
Bananas Posted June 13, 2014 Posted June 13, 2014 Just a random thought, try powering the switch connecting the server into the ups too.
Michael Posted June 13, 2014 Posted June 13, 2014 The server has a UPS I have done the switch it off for 5 mins bit and when I replaced the USB cable to the backup I did an internal inspection and it was clean. As I said, doesnt always happen at 1.20, in fact I think it may happen earlier as it did on Tuesday and it is only when the class logs on it becomes apparent. I knew earlier on Tuesday as a teacher happened to be doing something. The switches, of which they are 4, are DLinks but I can't say the model offhand. No weird events in event viewer and no SIMS or other MIS on my server (we have two networks). No odd devices either - they are all on the admin network. I like the idea of the share creation - will have a go at that - thanks Apologies I must of mis-read what you wrote regarding UPSes. If it's an APC UPS, open the Powerchute software and check for activity. If all appears normal, try doing a manual self test whilst physically in front of the server (not remotely), to see if the UPS is performing correctly. If you suspect something is wrong, then I'd advise you temporarily shutdown the server and run it without a UPS. You can then run other tests against the UPS using a basic workstation for example.
witch Posted June 13, 2014 Author Posted June 13, 2014 So, do you think it might be the UPS then? How so?
teejay Posted June 13, 2014 Posted June 13, 2014 Are there any firmware updates for your switches, that may be a problem. The other thing I can think of is use Wireshark to see if there is anything blitzing the server. Other than that, get your consultant in and gaffer tape him to the chair until they fix it. Drawing tacks may speed up the process in this case ;-)
Michael Posted June 14, 2014 Posted June 14, 2014 So, do you think it might be the UPS then? How so? Only speculation - but if the server logs look OK, it could be the UPS (when self testing) not working 100% correctly.
witch Posted June 14, 2014 Author Posted June 14, 2014 Are there any firmware updates for your switches, that may be a problem. The other thing I can think of is use Wireshark to see if there is anything blitzing the server. Other than that, get your consultant in and gaffer tape him to the chair until they fix it. Drawing tacks may speed up the process in this case ;-) He's great actually - basically because I am on my own and part-time the schools pay per year for phone support from this company and a techie for 2 hours a month. I have known him 12 years and he has taught me just about everything that I know - but he can only be with me when he is scheduled to be and this is such a pig of a fault that everyone is struggling. I don't know about firmware on the switches - will look into that
TheRobins Posted June 14, 2014 Posted June 14, 2014 He's great actually - basically because I am on my own and part-time the schools pay per year for phone support from this company and a techie for 2 hours a month. I have known him 12 years and he has taught me just about everything that I know - but he can only be with me when he is scheduled to be and this is such a pig of a fault that everyone is struggling. I don't know about firmware on the switches - will look into that Your position can be very frustrating, I have been there myself a few years ago. I'm not sure you need the highest level Microsoft certified bod in but just another pair of eyes and hands can make the world of difference. If staff are making enough noise, and say its causing them enough problems I am sure you can approach the business manager/head and request external support. Your issue could be warning signs of an imminent failure and the longer it is left the bigger the issues may be in the future. I suspect the DLINK switches may not be helping your cause but without the correct equipment or tools to do some more in depth tests you could end up going around in circles. Sadly doesn't look as if you are local else I would have offered to come over with some kit and spare switch and have a look, maybe there is someone local on here you may be able to give you a quick hand with it. Another little thought that just came to mind, How many NIC's does your server have? If you have more than one are they both being used? 1
witch Posted June 14, 2014 Author Posted June 14, 2014 Only one NIC I'm afraid What worries me is that unless we can actually catch the fault happening it is really difficult to know quite what to do next. I will try wireshark although I have always found it hard to interpret and I will look at the UPS and switch updates. If we got someone in it will cost a lot and there is no guarantee that they will find anything .
hallb15 Posted June 14, 2014 Posted June 14, 2014 This is a real puzzle! I agree it's hard to diagnose when you are only there a few hours at a time. It could be an IP address conflict with a device that has a static IP addr set, but not excluded in the DHCP scope. Or, has someone put a small hub on the network in an office somewhere and not told you? I have just bought some new switches so I could lend you the old ones which I know are working 100%. They're not smart or managed but at least you could rule them out if if doesn't fix the problem. I know how hard it is to manage on a non existent budget! PM me if you want to arrange something... 2
witch Posted June 14, 2014 Author Posted June 14, 2014 No hubs on my network. There are a couple of small switches in the school but if they were a problem then surely it would only affect the computers attached to them?
SpuffMonkey Posted June 14, 2014 Posted June 14, 2014 Had a similar problem to this years ago - caused by an electric kettle causing a surge - can you get any type of protection for the switches on even a temporary basis?
hallb15 Posted June 15, 2014 Posted June 15, 2014 No hubs on my network. There are a couple of small switches in the school but if they were a problem then surely it would only affect the computers attached to them? Not necessarily. Cheap consumer grade hubs and switches can spread bad packets like wildfire when they are overloaded or over heating. I once found a hub in the caretakers office at a school I used to work at that the CCTV company had kindly installed, inside a locked metal cupboard, so the caretaker could monitor the CCTV on his desktop computer. Whenever he inspected any of the recorded CCTV images, the whole school lost connectivity intermittently. Took ages to track that one down!
witch Posted June 15, 2014 Author Posted June 15, 2014 Ah Didn't know that Another thing to look at - how? No idea.... Yay!
Duke5A Posted June 16, 2014 Posted June 16, 2014 (edited) What model HP server and what kind of drives? I've got four DL180G5 servers in the datacenter used for CCTV recording and had nothing but trouble with them last year munching drives. Symptoms were similar to what your describing with the logs being clear and the filesystem randomly locking up. Given enough time individual drives would fail. It turns out a number of 750GB and 1TB drives had a faulty firmware on them. Check the models of each drive and get the latest firmware and do it for the controller as well. The only other thing I can think of is maybe jumbo frames got turned on somewhere. If a client/server is using it and one or the other/something in between doesn't support it then you'll see stuttering on the network links. This wouldn't account for the server physically locking up though - drive firmware would. Edited June 16, 2014 by Duke5A 1
CamelMan Posted June 16, 2014 Posted June 16, 2014 (edited) We have been having issues here where the server is fine but users suddenly cannot access work etc. I also noticed that I couldnt pull large files off and it would come up with "Insufficient Space" type messages (despite 10+G available) reboot and all was fine. It was our IRPStack (which can be knacked when installing/uninstalling things like antivirus software. Might be worth a look... How to fix "Not enough server storage is available to process this command" EDIT - We are on SVR2003 on the affected servers. Edited June 16, 2014 by CamelMan addition
witch Posted June 16, 2014 Author Posted June 16, 2014 Server fell over on Friday - HT couldn't even do Ctrl Alt Del but got an error message which I have left at work. When I got in this morning everyone descended on me - server looked OK until I tried to log on - when I got to the destop the icons stayed blank and when I tried to click on a shared drive - or the D drive on the server, it all greyed out on me. I had to reboot the server and all was well. I am going down the antivirus route at the moment and installing the updated version - you never know, that might be it. Goodness I hope so - I am slowly losing the will to live
psydii Posted June 16, 2014 Posted June 16, 2014 (edited) Might it be snapshot related? Edited June 16, 2014 by psydii
fiza Posted June 16, 2014 Posted June 16, 2014 Might it be snapshot related? I think @witch said it was not a virtualised server so no snapshots involved.
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now