3.11.8
2/16/2016
When a task error occurs, the task error code is supposed to turn off all IO. It should report the Task Error, important registers and stack data to the screen as well as save it in battery backed RAM for later retrieval. If a debugger is attached, the information is reported to the debugger also. Unless there is a debugger, all tasks are shutdown and the controller waits for a reset event. If there is a debugger, the debugger task is allowed to continue running and the task that caused and reports the task error is forced into an infinite sleep loop, allowing the cause of the TE to be investigated through the debugger.
However, over the passed couple of years, attempts to fix minor issues with this feature have caused other issues.
The Task Error Reporting feature was resetting in the middle of reporting task errors. From the users perspective this would look like a blue screen, with maybe a few numbers on the screen, followed by a quick reset. If they were not watching the screen at the moment of reset it would appear as if the machine stopped running on its own with no explanation.
Three main changes were made to resolve the issue:
- The Operating System watch dog is now being reset within the loop that writes out the Task Error Data. The loop was too long not to have the reset occur within it. Earlier versions of the code had this but it was inadvertently taken out somewhere along the way.
- There is a watchdog interrupt on the input task. If the input task does not reset the watch dog, the interrupt fires and initiates a task error. While resetting the OS watch dog the TE code is now also resetting the input task Watch Dog, which prevents a second TE from occuring while trying to report the first one.
- Interrupts now remain disabled until all of the TE data has been reported to the screen and saved to battery backed ram. This prevents any interrupts from modifying the data we want to preserve or more importantly, causing another Task Error. If a debugger is attached, interrupts are then enabled to allow the debug task to run and interact with the debugger.
A V4 controller came in from the field that was reporting large negative numbers, in the millions, for axis that have analog feedback. Forcing the memory to be cleared by removing the battery was the only thing that fixed it.
On the V4, the analog feedback comes from the DSP. Analog is read and stored in a 16 bit variable that has a numeric range of 32,767 to -32768. When it comes to the x86 it is stored into a 32 bit variable but it gets sign extended into the extra 16 bits. It was not possible to report values outside of the 16 bit range unless some other math was taking place befor the value was reported.
A code review revealed that the analog value had a 32 bit reference variable subtracted from it and that nothing in the code was setting the reference variable. Somehow the variable got set to a positive value larger than the 16 bit range and the subtraction resulted in a very large, in the millions, negative number.
This reference variable appears to be a hold over from a prior implementation. We don't expect to be making any other changes or new development in this code base so we chose the easiest fix. We zero out the reference variables when the PC sets the model.