Page 15 of 23
Re: OpenNetHomeServer
Posted: Tue Jul 26, 2016 5:43 pm
by Mastiff
I'm afraid that didn't help. I seem to get two to two and a half hour of time after a full restart (EG and NHS). So after 4-5 of the loss of contact things I changed to "localhost". That gave me a loss of contact after a few minutes. But I didn't restart NHS and EG completely, so I've done that now, and we'll see what happens. This is soooo annoying!
Edit: I tried to go up to a timeout for socket connections on 180 instead of 60 seconds. I see that I have 600 at the cabin, so that would be ten minutes. Maybe I have changed that down to 60 during the last week without remembering. There is just too much going on in my head right now, so the brain's turned to mush...
Re: OpenNetHomeServer
Posted: Tue Jul 26, 2016 5:59 pm
by Mastiff
180 didn't help. Going up to 600 now.
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 7:06 am
by Mastiff
That didn't help either. I also took a one month old copy of the VM out of the archive and tried running that, but it didn't help. The only thing I can think of that I need to try is to rename the VM. At the moment it's "Automatiseringsserver", which is of course too long for NetBIOS, but I can't imagine that would create socket errors when using an IP address to connect. My final attempt will be to remove the internal network and only run the external network (which is outside of the "physical" server, but inside the VM M0n0wall firewall that's running with the server as a host and dedicated network cards). It's so weird how this can happen after so long without any problems!
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 8:27 am
by krambriw
And I have no more (easy to implement ideas) either, in the python code I just create and use a socket connection according to the book but what happens then in the various networking layers is of course out of my control.
Are you using many events from ONH in your automation (I could only see RFX events in your log, maybe you are converting them)?
In my own setup here I am using a different kind of setup than you because I had similar network connection interrupts as you experience now. So it looks like this:
- ONH is running in a RaspberryPi
- a Python script is running in the same and publishing events to a MQTT broker also in the RPi
- in EventGhost in another machine, I use the MQTT Client to subscribe to the events
This has proven to be a very stable solution, I believe it is because the events from ONH are captured locally by the Python script instead of via the network. On the other hand, that is exactly how you have configured your system and that is still creating problems. And in my case, I just started a test yesterday and connected the EG plugin to the same ONH (in parallel to the Python script) and it is working perfect, no problems with a 60 sec time out value.
Considering that you have very little time left to spend on this, I wonder what you would be willing/capable of doing? If I just think about possible scenarios/modifications to try out I have some ideas.
BR Walter
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 8:46 am
by Mastiff
Well, at the moment I have been running for an hour or so after taking a few things in the house off line (I suspect it may be a misbehaving switch or something here). But of course that's before the trouble usually starts. Around 90 minutes to two hours is usually where the fun starts. And the reason that there's no Tellstick. (as I have them labeled) in the log is that the connection was lost before I fired up the log. Here's sometthing a bit more representative:
Code: Select all
10:39:18 RFXtrx.Type: Viking 02811 id: 6144 ' temperature: +6.6 deg C signal: 5 battery: 9'
10:39:22 RFXtrx.Type: Viking 02811 id: 7936 ' temperature: +22.9 deg C signal: 5 battery: 9'
10:39:22 Tellstick.FineOffset|1055 'Temp|22.9'
10:39:23 RFXtrx.Type: THGN122/123, THGN132, THGR122/228/238/268 id: 58881 ' temperature: +25.5 deg C humidity: 63 %RH status: normal signal: 6 battery: 9'
10:39:24 RFXtrx.Type: Viking 02811 id: 9984 ' temperature: +31.3 deg C signal: 4 battery: 9'
10:39:27 RFXtrx.Type: Viking 02811 id: 32256 ' temperature: +19.9 deg C signal: 5 battery: 9'
10:39:27 Tellstick.FineOffset|1150 'Temp|19.9'
10:39:28 RFXtrx.Type: Viking 02811 id: 59904 ' temperature: +17.3 deg C signal: 4 battery: 9'
10:39:28 RFXtrx.Type: Viking 02811 id: 13312 ' temperature: +19.8 deg C signal: 5 battery: 9'
10:39:31 RFXtrx.Type: Viking 02811 id: 57088 ' temperature: +21.9 deg C signal: 5 battery: 9'
10:39:32 RFXtrx.Type: Viking 02811 id: 45056 ' temperature: +20.1 deg C signal: 5 battery: 9'
10:39:38 RFXtrx.Type: Viking 02811 id: 32768 ' temperature: +19.6 deg C signal: 5 battery: 9'
10:39:38 RFXtrx.Type: Viking 02811 id: 58624 ' temperature: +23.0 deg C signal: 5 battery: 9'
10:39:38 Tellstick.FineOffset|1253 'Temp|23.0'
10:39:41 RFXtrx.Type: Viking 02811 id: 43776 ' temperature: +21.6 deg C signal: 5 battery: 9'
10:39:46 RFXtrx.Type: Viking 02811 id: 6400 ' temperature: +21.5 deg C signal: 5 battery: 9'
10:39:46 Tellstick.FineOffset|1049 'Temp|21.5'
10:39:47 RFXtrx.Type: Viking 02811 id: 2048 ' temperature: +22.3 deg C signal: 4 battery: 9'
10:39:52 RFXtrx.Type: Viking 02811 id: 60160 ' temperature: +18.4 deg C signal: 4 battery: 9'
10:39:52 Tellstick.FineOffset|1259 'Temp|18.4'
10:39:56 RFXtrx.Type: Viking 02811 id: 3328 ' temperature: +22.2 deg C signal: 5 battery: 9'
10:40:02 RFXtrx.Type: THGN122/123, THGN132, THGR122/228/238/268 id: 58881 ' temperature: +25.5 deg C humidity: 63 %RH status: normal signal: 6 battery: 9'
10:40:02 Tellstick.Oregon|110|1| 'Temp|25.5|Hum|63|Batt|0|'
10:40:02 RFXtrx.Type: AC address: 008d303a unit: 08 command: off ' level: 0 signal: 5'
10:40:06 RFXtrx.Type: Viking 02811 id: 28416 ' temperature: +20.9 deg C signal: 5 battery: 9'
10:40:06 RFXtrx.Type: Viking 02811 id: 6144 ' temperature: +6.6 deg C signal: 5 battery: 9'
10:40:10 RFXtrx.Type: Viking 02811 id: 7936 ' temperature: +22.9 deg C signal: 5 battery: 9'
10:40:10 Tellstick.FineOffset|1055 'Temp|22.9'
And you're right, I don't have any time left. I can't introduce new stuff here now, I just have to find out what the hell is going on, or find a nice way around it. The thing is that the Tellstick is most of all a backup solution. I send every event (oven on/off, amp on/off) both through the RFXtrx and the Tellstick, to make sure that at least one of them will take, and more important, so that if one of the devices fail there will still be control. I must try to work around this. Up to now I have been restarting ONH (a batch file that first kills the service, then uses two 5000 ms pings to a non existent network device to have a 10 second wait and then restart ONH) and at the same time disabling the plug-in and re-enabling it four seconds later. That's actually the biggest problem because it locks up EG for 10 seconds, which means that if you try to turn on an amp or something when that happens, it looks like there's no reaction. So I have changed the macro to only kill and restart ONH, and I will see if it's able to connect to the plug-in again without a reset of the plugin.
Oh, and there are less events going into the Tellstick/ONH because I have a large antenna on the RFXtrx, but a smaller one on the Tellstick.
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 8:48 am
by krambriw
I think I have a solution!
I was playing with the settings and lowered the time out setting to 5 seconds, and yes I got a lot of socket errors. But funny enough, I could see that the connection was still working since events came in anyway. So I thought, let's try to comment out the whole error handling and see what is happens. And it works fine, no need to make a re-connection!
The setting should still be reasonable high because it will determine the time out for the 'TellStick Duo is disconnected' event.
What is the drawback? How is a "real" disconnection repaired? That I cannot answer for the moment...needs to be tested
What you can do is to modify the plugin code (you find it around line 1325) like this:
Code: Select all
except socket.error, e:
pass
# print 'socket.error', self.atStart, e
# self.connectionError = True
# reConnects += 1
# #if reConnects > 4:
# if reConnects > 1:
# if self.atStart == 1:
# eg.PrintError(self.text.tcp_connection_error)
# eg.TriggerEvent(
# self.text.tcp_connection_error,
# prefix = self.prefix
# )
# self.atStart = 2
# self.keepAliveThreadEvent.set()
# reConnects = 0
# if self.atStart == 0:
# eg.PrintError(self.text.tcp_connection_at_start)
# eg.TriggerEvent(
# self.text.tcp_connection_at_start,
# prefix = self.prefix
# )
# self.atStart = 3
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 9:00 am
by Mastiff
Fantastic, maestro!

I just have to see if the removal of the stuff from the network changes anything. I still do not have any errors, but I will do this change when (or hopefully if...) I get one again, and then I will report back to tell you what happens. Thanks a lot for helping me out here!
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 9:03 am
by Mastiff
Stefan came in with an answer now, and he thinks it's the plug-in that disconnects for some reason:
http://forum.opennethome.org/viewtopic. ... p=648#p648
That would probably mean that you're really on to something with this solution, right?
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 9:15 am
by krambriw
I did a test and disconnected the network cable to the ONH for some minutes. When I connected it again, events started to come in!!! So it is still working even after a network break. I guess this will do. If you want to monitor the connection to ONH (maybe not needed for you since everything is in the same VM?) you would have to build something in python in EG that is checking that ONH events are coming in at a regular interval or so.
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 9:40 am
by Mastiff
Brilliant, Walter!

Then I will just comment out this stuff and get it back working. Thanks a lot! As for checking, that's really easy: A reaccuring timer that resets and starts again every time any Tellstick event comes in.

Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 11:03 am
by krambriw
Stefan came in with an answer now, and he thinks it's the plug-in that disconnects for some reason
Yes. that is correct, the plugin will always re-connect when the socket error is triggered (but that code is now commented out). Obviously, in your case, this is some kind of "false alarm" and I'm not 100% happy with this fix, how should a "real" socket error then be handled? Anyway, I will go for this solution and leave it to EG to monitor (with a timer solution as you mention) and eventual routine to restart the plugin
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 12:16 pm
by Mastiff
How about this for a workaround (remember I'm an idiot when it comes to programming, so I may just be embarassing myself here): When the socket error comes in, it triggers a timer. If there are no actual events coming in from the NHS instance within a set time (one minute would be more than enough for me), then it triggers the reset code. Does that make sense?
Oh btw, I may actually have found the source of my particular problem. I have now been running flawlessly for almost three hours, which hasn't happened since this started. I do believe it was the 8 port gigabit switch that's been running solidly for 4-5 years that suddenly found out that it wanted to f*ck with me now that I'm selling the house and only keeping a position as sys admin on remote call.
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 12:32 pm
by krambriw
Your idea makes sense! Good input!
I'll be able to implement and test this myself as well, I can easily prove a connection break

I guess for this we have time if the system is running ok now, you might wan't to update the system later with an updated plugin once finalized
Oh btw, I may actually have found the source of my particular problem
Those buddies does NOT like the thought of a new master

Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 2:38 pm
by Mastiff
Wow, I finally say something that makes sense! Haven't done that in a few weeks...

Unfortunately something weird happened here (and I haven't implemented the code commenting out stuff that you gave me): Suddenly, a couple of hours ago it seems the NHS plug-in stopped sending events to the logger. But it's still alive. No error messages, and I can see that it's still doing actions (turning on and off stuff) because I get events like this from the RFXtrx:
Code: Select all
16:33:52 RFXtrx.Type: AC address: 008d303b unit: 01 command: off ' level: 0 signal: 6'
I tried to "Restart all" in the plug-in dialogue, but nothing happened. Disdabling the plug-in didn't change anything either. Final attempt was to restart NHS, and that fixed it. Any idea what that can have been?
Re: OpenNetHomeServer
Posted: Wed Jul 27, 2016 4:00 pm
by krambriw
I think the handling of threads etc (those you can see listed in the plugin list) could be improved, especially handling of those buttons. Sometimes it is easier to just select the logger action and execute (run) it again. Since we have stopped & started a bit now and then, it might be that things got out of sync. But as long as a fresh restart is working as expected, we should be fine