15 KiB
Modbus RTU Monitor Integration - Development Guide
Overview
This integration provides passive monitoring of Modbus RTU communication over TCP connections. It listens to existing Modbus traffic between a gateway and slave devices (e.g., thermostats) and creates Home Assistant entities based on discovered data.
Key Feature: Non-intrusive monitoring - the integration does not initiate Modbus requests, it only observes traffic. The exception is the climate entity's setpoint write functionality.
Use Case & Background
This integration was developed to enable remote control and monitoring of proprietary thermostats that are integrated into a complex PLC-based heat automation system via Modbus RTU.
Real-World Scenario
Current Setup:
- Multiple proprietary thermostats connected via 2 Modbus RTU lines
- Complex PLC (Programmable Logic Controller) manages heat automation
- PLC writes configuration registers to thermostats (time, mode, pump status)
- PLC reads sensor data from thermostats (temperature, humidity, setpoint)
Challenges:
- Proprietary thermostats with no native Home Assistant support
- Cannot interrupt or interfere with critical PLC communication
- Need to read temperature/humidity and set setpoint remotely
- Existing automation must continue functioning during transition
Solution Approach:
- Passive monitoring: Observe PLC ↔ thermostat communication without disruption
- Selective writes: Only write setpoint when user changes it in HA
- Gradual migration path: Enables Home Assistant integration while PLC continues operating
- Future goal: Eventually replace PLC entirely with Home Assistant automations
Why Passive Monitoring Matters
The PLC continuously communicates with thermostats to:
- Synchronize time (hour/minute registers)
- Set operating modes (summer/winter mode)
- Control heat pump status (zone pump register 211)
- Read temperature, humidity, and manual mode setpoint
Critical: Any polling or active queries from Home Assistant could:
- Disrupt time-sensitive PLC communication
- Cause Modbus bus contention
- Potentially confuse the PLC or thermostats
- Create unpredictable heating behavior
Passive monitoring solves this by only observing existing traffic, making the integration completely invisible to the PLC and thermostats (except for explicit setpoint writes initiated by users).
Migration Strategy
-
Phase 1 (Current): Passive monitoring + manual setpoint control
- Read all sensor data (temp, humidity, registers)
- Allow manual temperature adjustment from HA
- PLC continues all automation logic
-
Phase 2 (Future): Gradual automation takeover
- Start replacing PLC automations with HA automations
- Monitor additional registers to understand PLC logic
- Test HA automations alongside PLC
-
Phase 3 (Goal): Full HA control
- HA handles all heating automation
- PLC decommissioned or relegated to backup role
- Complete integration with HA ecosystem
This integration is the foundation for this migration path.
Architecture
Core Components
-
hub.py - Central communication hub
- Manages TCP connection to Modbus gateway
- Passively decodes Modbus RTU frames
- Handles reconnection logic for extended outages
- Provides active write capability for setpoint changes
-
coordinator.py - Data update coordinator
- Manages entity updates efficiently
- Uses throttled updates (5-second minimum interval)
- Coordinates between hub and entity platforms
-
Entity Platforms:
- climate.py - Thermostat control with HVAC action display
- sensor.py - Humidity and register monitoring
- binary_sensor.py - Coil states and zone pump status
-
const.py - Constants and data models
- Register addresses and ranges
- SlaveData and ModbusRTUFrame dataclasses
- User-friendly register name mappings
Critical Implementation Details
Reconnection Logic (CRITICAL)
The integration must handle extended server outages (1-4+ hours) gracefully:
Key Requirements:
- Monitor task is always created - even if initial connection fails (hub.py lines 198-220)
- Always null
_readerand_writerafter failed reconnection attempts - Check
if self._reader is None:before attempting to read - Exception handler must detect connection errors and trigger reconnection
- Buffer must be cleared after reconnection attempts
- Reconnection delay is 30 seconds (prevents DDOS during outages)
Files: hub.py lines 198-333
Why This Matters:
- Without the monitor task being always created, integrations loaded while server is down never attempt reconnection
- Without proper state management, the integration enters an infinite error loop after the first failed reconnection
- Both scenarios require manual reload without these fixes
Critical Design Decision:
- The monitor task handles ALL connection logic (hub.py:254-285)
async_setup()only creates tasks, never attempts connection directly (hub.py:198-212)- This prevents race conditions between setup and reconnection logic
- First connection attempt is immediate (no delay)
- Subsequent reconnection attempts use 30-second delay to prevent DDOS
Passive Monitoring Pattern
The integration monitors three types of Modbus traffic:
-
Temperature/Humidity (Register 0x83, 4 registers)
- Detected when request matches TEMP_HUMIDITY_REGISTER
- Response contains temperature (index 2) and humidity (index 3)
- Values divided by 10 for actual reading
-
Coils (Function code 0x01, 40 coils starting at address 1)
- Binary states (ON/OFF)
- Disabled by default (user can enable)
-
Register Ranges (Function code 0x03):
- Range 1: 165-184 (20 registers) - REGISTER_START_ADDRESS
- Range 2: 210-217 (8 registers) - REGISTER_START_ADDRESS_2
- Range 3: 140-162 (23 registers) - REGISTER_START_ADDRESS_3
Pattern: Hub detects request frames, stores pending request data, then matches response frames to extract values.
Register Name Mapping
User-friendly names are maintained in const.py REGISTER_NAMES dictionary:
REGISTER_NAMES: dict[int, str] = {
144: "Setpoint", # Target temperature (written by PLC or HA)
145: "Max Setpoint", # Maximum allowed temperature
146: "Min Setpoint", # Minimum allowed temperature
154: "Hour", # Current hour (synchronized by PLC)
155: "Minute", # Current minute (synchronized by PLC)
156: "Day of week", # Current day (synchronized by PLC)
157: "Current temperature", # Actual temperature reading
179: "Outside temperature", # External temperature sensor
211: "Zone pump", # Heat pump/zone valve status (1=ON, 0=OFF)
}
PLC-Written Registers: The PLC writes registers 154-156 (time sync) and 211 (pump control). These are monitored to understand system state but should not be written by HA while the PLC is active.
HA-Writable Register: Register 144 (Setpoint) can be safely written by Home Assistant when users adjust temperature, as this mimics manual thermostat adjustment.
To add new mappings: Update REGISTER_NAMES in const.py - sensor entities will automatically use friendly names. As you discover more registers through monitoring, document their purpose here.
Climate Entity Bidirectional Sync
The climate entity maintains setpoint synchronization in both directions:
-
User sets temperature in HA:
- Climate entity writes to device via hub.async_write_setpoint()
- Immediately updates coordinator data to prevent UI flicker
- Scaled value (temp × 10) written to register 144
-
Physical thermostat changes setpoint:
- Hub passively monitors register 144
- Climate entity reads from coordinator.data[slave_id].registers[144]
- Scales value (÷ 10) for display
Key Detail: Register 144 value is immediately updated after write to prevent UI from showing stale value until next monitoring cycle.
HVAC Action Display
Climate entity shows heating status based on register 211:
HVACAction.HEATINGwhen register 211 = 1HVACAction.IDLEwhen register 211 = 0 or unavailable
Dual Representation: Register 211 exists as both:
- Regular sensor showing raw integer value (disabled by default)
- Binary sensor "Zone pump" showing ON/OFF state (enabled by default)
This allows users to choose their preferred representation.
Entity Default States
- Enabled by default: Humidity sensor, zone pump binary sensor, climate entity
- Disabled by default: All register sensors (except humidity), all coil sensors
Users can selectively enable the entities they need.
Data Flow
Modbus Gateway ←→ TCP Connection ←→ Hub (passive monitoring)
↓ (frames decoded)
discovered_slaves dict
↓ (throttled updates)
Coordinator
↓
Entity Platforms (climate, sensor, binary_sensor)
Update Throttling
Coordinator Updates: Minimum 5-second interval between updates
- Defined in COORDINATOR_UPDATE_INTERVAL
- Prevents overwhelming Home Assistant with high-frequency Modbus traffic
- Hub calls
_update_coordinator_throttled()which checks timing
Why: Modbus traffic can be very frequent (multiple times per second). Without throttling, HA UI becomes sluggish.
Connection Handling
Multiple Servers
Users can configure multiple Modbus gateways as separate config entries. Each entry creates its own:
- Hub instance with dedicated TCP connection
- Monitor task
- Set of discovered slaves and entities
Important: 30-second reconnection delay prevents DDOS when multiple servers are offline simultaneously.
Availability Tracking
Slaves are marked unavailable if no data received for 5 minutes (AVAILABILITY_TIMEOUT).
Implementation: Separate _availability_check_loop() runs every 60 seconds and compares last_seen timestamp.
Modbus Protocol Details
RTU Frame Structure
[Slave ID][Function Code][Data...][CRC Low][CRC High]
CRC Calculation: Uses Modbus RTU CRC-16 algorithm (polynomial 0xA001)
Supported Function Codes
- 0x01: Read Coils (request + response with bitmap)
- 0x03: Read Holding Registers (request + response with values)
- 0x06: Write Single Register (used for setpoint writes)
Signed vs Unsigned Integers
Modbus uses unsigned 16-bit values, but physical readings can be negative (e.g., temperatures).
Conversion (hub.py):
if value > 32767:
signed_value = value - 65536
else:
signed_value = value
This converts unsigned to signed int16 representation.
Common Development Tasks
Adding New Register Mappings
- Identify register address and purpose through monitoring
- Update
REGISTER_NAMESin const.py:REGISTER_NAMES: dict[int, str] = { # ... existing mappings ... XXX: "Your Register Name", } - Sensors automatically display friendly name
Extending Monitored Register Ranges
To monitor additional register ranges:
-
Add constants to const.py:
REGISTER_START_ADDRESS_4 = 0xYY # Your start address REGISTER_MONITOR_COUNT_4 = ZZ # Number of registers -
Update hub.py
_handle_frame()to detect requests/responses for the new range- Follow pattern from existing ranges (1, 2, 3)
- Use unique pending request key (e.g.,
f"{slave_id}_registers4")
-
Update sensor.py
async_setup_entry()to create sensors for new range- Add tracking set (e.g.,
added_slaves_registers4) - Add sensor creation logic in
_async_add_sensor_entities()
- Add tracking set (e.g.,
Testing Reconnection Logic
Critical Test: Long outage scenario
# 1. Start integration with working server
# 2. Stop Modbus server
# 3. Wait 2+ hours
# 4. Check logs - should see reconnection attempts every ~30 seconds
# 5. Start Modbus server
# 6. Verify automatic reconnection within 30 seconds
# 7. Verify data flow resumes without manual reload
Known Issues and Solutions
Issue: UI Flicker When Setting Temperature
Symptom: After setting temperature, UI briefly reverts to old value before updating.
Solution: Immediately update slave_data.registers[SETPOINT_REGISTER] after successful write (climate.py line 178).
Issue: Integration Won't Reconnect After Long Outage
Symptom: After server outage >1 hour, integration requires manual reload.
Root Cause: Reader/writer not properly nulled on failed reconnection.
Solution: Implemented in hub.py lines 222-333 (see Reconnection Logic section above).
Issue: Too Many Connection Attempts During Outage
Symptom: Server logs show connection spam when offline.
Solution: 30-second reconnection delay (hub.py line 237).
Code Patterns and Conventions
Async Safety
- All I/O operations are async
- Write operations use
self._write_lockto prevent concurrent access - Coordinator updates use
@callbackdecorator where appropriate
Logging
- Debug level: Frame processing details, non-critical events
- Info level: Connection status, slave discovery, successful operations
- Warning level: Connection loss, reconnection attempts
- Error level: Failed reconnections, unexpected errors
Format: Always include self.host:self.port in connection-related messages.
Entity Registration
Entities are dynamically created when slaves are discovered:
- Use tracking sets (e.g.,
added_slaves,added_slaves_registers) - Subscribe to coordinator updates via
async_add_listener() - Only create entities when new slaves/data detected
Device Info
All entities for a slave share the same device:
DeviceInfo(
identifiers={(DOMAIN, f"{coordinator.config_entry.entry_id}_{slave_id}")},
name=f"Modbus Slave {slave_id}",
manufacturer=MANUFACTURER,
model=MODEL,
)
Configuration
No YAML configuration - integration uses config flow only.
Config Entry Data:
host: Modbus gateway IP/hostnameport: TCP port (default 502)
Runtime Data: Hub instance stored in entry.runtime_data
Files Overview
- manifest.json - Integration metadata and dependencies
- const.py - Constants, register addresses, data models
- hub.py - TCP connection, frame decoding, reconnection logic
- coordinator.py - Data update coordination
- config_flow.py - UI configuration
- climate.py - Thermostat entity
- sensor.py - Humidity and register sensors
- binary_sensor.py - Coil and zone pump sensors
- __init__.py - Integration setup/unload
Important Notes for AI Assistants
- Never remove the 30-second reconnection delay - it prevents DDOS during outages
- Always clear buffer after reconnection - prevents processing stale frames
- Maintain throttling - Modbus traffic is high-frequency, HA needs protection
- Preserve passive monitoring - don't add active polling unless specifically requested
- Keep register sensors disabled by default - there are many, most users don't need them
- Test with extended outages - this is the most common production issue
Future Enhancements
Potential improvements to consider:
- Support for Modbus RTU over serial (currently TCP only)
- Configurable register ranges via UI
- Entity auto-discovery based on register schema files
- Diagnostic sensors for connection statistics
- Service to manually trigger reconnection