bandsaunter/bandsaunter/acurite.py
The Dust Council 872eadac37 Read the parity the way the sensors write it
Their sensors are on the air, five of them, and every message was arriving
intact.  The bytes they sent, recovered from their own capture:

    A7 1B 44 A6 09 4B 00     channel B  id 271B   22.7 C  38%
    F9 35 44 2B 09 C9 6F     channel A  id 3935   22.5 C  43%
    21 E5 84 1E 0A 9F 51     channel C  id 21E5   31.1 C  30%  battery low
    F9 7D 44 28 09 50 3B     channel A  id 397D   23.2 C  40%
    0D D6 44 A9 09 CF A8     channel C  id 0DD6   23.1 C  41%

Every checksum correct, every message type 0x04, every reading plausible.  The
framing was right, the bit offset was right, the byte order was right, the
pulses had been recovered perfectly for days.  One bit of convention was
wrong: the parity in the top bit of each payload byte is even, and this
required it to be odd.  Twenty payload bytes across five independent messages,
every one of them even, which is not something twenty bytes do by chance.

That is the whole fault.  Everything else changed in this and the two commits
before it was real and worth doing, and none of it was why nothing decoded.

Three things follow.

The five messages are now a test, checked byte for byte against the weather
they carry.  They are worth more than everything else in that file put
together: every other test there puts a reading in through an encoder written
from the same description as the decoder, so the two agree by construction and
agree about anything they are both wrong about -- which is exactly what
happened.  An encoder tested against its own decoder cannot find a fault in
the description they share, and no amount of it would ever have found this.

The emptiness check earns its place now.  Odd parity rejects a byte of all
zeroes; even parity accepts one, so a run of silence read as zeroes satisfies
both the parity and a sum of zero, and the only thing standing between that
and a display full of sensors is the test that some byte is non-zero.  It was
there for tidiness and is now load-bearing; the comment says so.

And the readings of a burst are tried in order and the search stops at the
first that yields anything, rather than pooling them.  Half a dozen readings
at two byte orders is sixteen times the chances for a coincidence to satisfy a
twelve-bit check, and sensors that were not there began appearing in the
invented garden the moment the alternatives went in -- caught by the test that
asks whether everything heard is something that exists.  Stopping early costs
nothing: a burst that reads correctly the ordinary way never reaches the
alternatives, and one that does not reaches them exactly as before.

Full suite 2351 passed, checked against three more deliberately broken builds.
Sixty seconds of receiver noise yields nothing and eight hundred seconds of
the invented garden yields no sensor that is not there.  Built as
2026-09-07_05.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PsWPTweCT6pwxKngvVxcg
2026-09-07 23:03:16 -07:00

1914 lines
85 KiB
Python

"""AcuRite weather sensors: the messages they send, and how to get them off
the air.
A consumer weather station is two things. The display on the kitchen wall is
one of them; the other is a plastic box on a fence post that says what it can
see every sixteen seconds or so, in the clear, on 433.92 MHz, to anyone who
happens to be listening. This reads the box.
**What is here.** Five families, each with its own framing and its own check:
=========================== ===== ==================================
model bytes what it says
=========================== ===== ==================================
Tower 592TXR / 06002RM 7 temperature, humidity
5-in-1 06014RM / VN1TXC 8 wind, direction, rain / wind, temperature, humidity
Lightning 6045M 9 temperature, humidity, strikes, how far off
609TXC 5 temperature, humidity
606TX 4 temperature
=========================== ===== ==================================
The first three share a framing -- two bytes of identity, a byte saying what
kind of message this is, the payload, and a checksum which is the sum of
everything before it -- and the payload bytes each carry even parity in their
top bit. That is between twelve and fourteen bits of check on a message of
seven to nine bytes, which is enough to accept a message on a band where
doorbells, car keys, tyre sensors and someone's garage all transmit.
The last two are older and cheaper and carry one byte of check between them:
a sum on the 609 and a CRC-8 on the 606. Eight bits is one false message in
two hundred and fifty-six tries, and this looks at every bit offset of every
burst, so eight bits on its own is not enough. Those two are therefore only
believed when the *same* message arrives twice in one burst -- which costs
nothing, because these sensors send every message three times in a row for
exactly this reason.
**What is not here.** The Atlas, the 986 and 515 fridge thermometers, the
00275rm room monitor and the 899 standalone rain gauge. They are on the same
band and are not read. A message from one of them whose framing happens to
match is reported as an unknown message type with its identity and nothing
else, rather than guessed at.
That last part is not a nicety. The Atlas is nine bytes like the lightning
detector and lays its payload out differently, and nothing but the message
type tells them apart -- so every decoder here insists on a type it knows,
and a nine-byte message that is not a 6045M is reported as a sensor of an
unknown model rather than as a lightning detector reading a temperature off
the wrong bits. Wrong weather under somebody's sensor name is a worse
answer than no weather at all.
**Where the numbers come from.** These formats are implemented from their
published descriptions and are checked here against frames built from the
same descriptions. That proves the framing, the parity, the checksums and
the arithmetic; it is not the same as having held every one of these
sensors. The reason it is nonetheless safe to run is that nothing is
reported which has not satisfied its own check.
**Units.** Different models report in different units -- the tower sends
tenths of a degree Celsius, the 5-in-1 and the lightning detector send tenths
of a degree Fahrenheit, and the rain gauge sends hundredths of an inch. All
of it is converted here, once, to Celsius, kilometres an hour, millimetres
and kilometres, so that a log holds one kind of number and the person reading
it can be shown whichever they prefer. The raw counts survive alongside the
converted values where they mean something on their own, which for the rain
gauge they do: it is a tipping bucket and the count only ever goes up.
"""
from __future__ import annotations
import math
from dataclasses import dataclass, field
import numpy as np
__all__ = ["Reading", "Measure", "format_measure",
"decode", "decode_burst", "candidates", "confirmed",
"readings_from", "MODELS", "MODEL_NAMES", "CHANNELS", "ACURITE_HZ",
"tower_frame", "five_in_one_wind_rain", "five_in_one_weather",
"lightning_frame", "frame_609", "frame_606",
"bursts", "Burst", "bits_pwm", "bits_ppm", "slicings", "pulse_train",
"modulate", "baseband", "WIND_POINTS", "compass", "parity8",
"crc8", "VirtualSensor", "SimulatedSensors", "default_sensors",
"survey", "Survey", "timings", "near_misses", "NearMiss",
"BIT_ORDERS"]
# Where every one of these sensors transmits. Nominally 433.92 MHz; the
# cheap transmitters in them drift by tens of kilohertz with temperature,
# which is why what follows works on the envelope of a band rather than on a
# carrier at an exact frequency.
ACURITE_HZ = 433_920_000.0
# ---------------------------------------------------------------------------
# The small arithmetic every one of these depends on
# ---------------------------------------------------------------------------
def parity8(value: int) -> int:
"""Odd-parity bit of one byte: 1 when an odd number of bits are set."""
value ^= value >> 4
value ^= value >> 2
value ^= value >> 1
return value & 1
def crc8(data, poly: int = 0x07, init: int = 0x00) -> int:
"""The plain byte-at-a-time CRC-8 the 606TX carries."""
crc = init
for byte in data:
crc ^= byte
for _ in range(8):
crc = ((crc << 1) ^ poly) & 0xFF if crc & 0x80 else (crc << 1) & 0xFF
return crc
# The channel switch on the back of a tower sensor has three positions, and
# they are not encoded in the order anyone would guess: A is 3, B is 2 and C
# is 0. Code 1 is not a switch position at all; the published descriptions
# call it E and so does this, rather than pretending it is a D.
CHANNELS = ("C", "E", "B", "A")
# The sixteen wind directions, in the order the sensor numbers them, which is
# also not the order anyone would guess. Index it with the low nibble.
WIND_POINTS = (315.0, 247.5, 292.5, 270.0, 337.5, 225.0, 0.0, 202.5,
67.5, 135.0, 90.0, 112.5, 45.0, 157.5, 22.5, 180.0)
_COMPASS = ("N", "NNE", "NE", "ENE", "E", "ESE", "SE", "SSE",
"S", "SSW", "SW", "WSW", "W", "WNW", "NW", "NNW")
def compass(degrees: float) -> str:
"""The sixteen-point name for a bearing, for people rather than plots."""
return _COMPASS[int((float(degrees) % 360.0) / 22.5 + 0.5) % 16]
# One tip of the bucket, in millimetres. The gauge counts hundredths of an
# inch, which is 0.254 mm, and the count is cumulative for the life of the
# battery.
RAIN_PER_COUNT_MM = 0.254
# What a sensor can physically report. A checksum can be satisfied by a
# message the hardware could not have sent, and on a band this crowded that
# happens often enough to be worth catching.
TEMPERATURE_RANGE = (-40.0, 70.0) # Celsius
HUMIDITY_RANGE = (0, 100) # per cent
WIND_RANGE = (0.0, 260.0) # km/h; the sensor tops out below this
# The message type a 6045M sends. Insisted on rather than assumed from the
# length; see :func:`decode_lightning` for why.
LIGHTNING_MESSAGE = 0x2F
# Every message type any decoder here recognises, for saying whether a burst
# that framed correctly is one of them.
_KNOWN_TYPES = frozenset({0x04, 0x31, 0x38, LIGHTNING_MESSAGE})
# Where in a burst a message is allowed to sit. A real one begins after the
# sync and runs to the end, so a handful of bits of sync in front of it and
# almost nothing behind. See :func:`candidates`, where these do the work.
LEAD_BITS = 8 # room for twice the four sync pulses these sensors send
TAIL_BITS = 4 # room for a couple of stray edges at the end
# A message this cannot read is held to the tighter figures: it has no type
# field that means anything here and no range that can be checked, so where
# it sits is the only evidence there is that it is a message at all.
UNREAD_LEAD_BITS = 4
UNREAD_TAIL_BITS = 1
def _fahrenheit(f: float) -> float:
return (f - 32.0) * 5.0 / 9.0
# ---------------------------------------------------------------------------
# What one message came to
# ---------------------------------------------------------------------------
@dataclass(frozen=True)
class Measure:
"""One quantity a sensor reported, with the unit it is held in."""
name: str
value: float
unit: str = ""
# What the sensor actually sent, where that is a different number from
# the one above: the rain counter, the strike distance in miles. Kept
# because it is the evidence and the converted value is a restatement.
raw: float | None = None
@property
def text(self) -> str:
return format_measure(self)
def format_measure(measure: Measure, imperial: bool = False) -> str:
"""One measurement as a person would write it, in either system."""
name, value, unit = measure.name, measure.value, measure.unit
if unit == "C":
return f"{value * 9 / 5 + 32:.1f} F" if imperial else f"{value:.1f} C"
if unit == "%":
return f"{value:.0f}%"
if unit == "km/h":
return f"{value / 1.609344:.1f} mph" if imperial else f"{value:.1f} km/h"
if unit == "mm":
return f"{value / 25.4:.2f} in" if imperial else f"{value:.1f} mm"
if unit == "km":
return f"{value / 1.609344:.0f} mi" if imperial else f"{value:.0f} km"
if unit == "deg":
return f"{value:.0f}° {compass(value)}"
if unit == "lux":
return f"{value:.0f} lux"
if unit:
return f"{value:g} {unit}"
# A bare count -- strikes, bucket tips -- reads better without a ".0".
return f"{value:g}" if value != int(value) else f"{int(value)}"
@dataclass
class Reading:
"""One message that satisfied its own checks, and what it said.
The bits are kept because they are the evidence; everything else in here
is an opinion about them, and the two should not be confused.
"""
model: str = "" # what to call it to a person
family: str = "" # what to file it under: tower, 5n1, ...
sensor: str = "" # the identity in the message, as hex
channel: str = "" # A, B or C, where the model has a switch
battery_low: bool | None = None
message: int = 0 # the message-type field, as sent
measures: tuple = ()
bits: str = ""
checks: tuple = ()
at: float = 0.0 # when it arrived, as a clock time
copies: int = 1 # how many times it arrived, identically
offset: int = 0 # where in the burst its first bit was
@property
def key(self) -> str:
"""What a friendly name is attached to.
Family and identity, not channel: the channel switch is on the
outside of the box and someone will move it, and moving it should not
lose the name they gave the box.
"""
return f"{self.family}/{self.sensor}"
def value(self, name: str):
"""One measurement by name, or None."""
for measure in self.measures:
if measure.name == name:
return measure.value
return None
def describe(self, imperial: bool = False) -> str:
head = f"{self.model} {self.sensor}"
if self.channel:
head += f" ch {self.channel}"
rest = " ".join(f"{m.name} {format_measure(m, imperial)}"
for m in self.measures)
if self.battery_low:
rest += " battery low"
if not self.measures:
rest = rest or f"message type {self.message:#04x}, not understood"
return f"{head} {rest}".strip()
# Every family, in the order they are tried, with how many copies of a
# message must arrive in one burst before it is believed. See the module
# docstring: it is a function of how much check the message carries.
MODELS = (
("tower", "Tower 592TXR", 7, 1),
("5n1", "5-in-1 06014RM", 8, 1),
("6045", "Lightning 6045M", 9, 1),
("609", "609TXC", 5, 2),
("606", "606TX", 4, 2),
)
MODEL_NAMES = {family: name for family, name, _bytes, _r in MODELS}
MESSAGE_BYTES = {family: n for family, _name, n, _r in MODELS}
MESSAGE_FAMILY = {n: family for family, _name, n, _r in MODELS}
REPEATS_NEEDED = {family: repeats for family, _name, _bytes, repeats in MODELS}
# ---------------------------------------------------------------------------
# The three that share a framing: tower, 5-in-1, lightning detector
# ---------------------------------------------------------------------------
def _txr_framed(data: list[int]) -> bool:
"""The checks the whole TXR family carries: a sum, and even parity.
The sum covers every byte before it. The parity is in the top bit of
each payload byte -- everything but the two identity bytes and the
checksum itself -- and is *even*: the bit is set so that the number of
bits in the byte comes out even.
Even rather than odd, which is worth saying plainly because this had it
the other way round and nothing whatever was decoded for it. Five
messages off five real sensors settled it: twenty payload bytes, every
one of them even, which is not something twenty bytes do by chance.
One consequence has to be handled here rather than assumed away. Odd
parity rejects a byte of all zeroes and even parity accepts it, so a run
of silence read as zeroes satisfies both this and a sum of zero -- which
is why the emptiness test below is a check and not a nicety.
"""
if (sum(data[:-1]) & 0xFF) != data[-1]:
return False
if not any(data):
return False
return all(parity8(byte) == 0 for byte in data[2:-1])
def _txr_identity(data: list[int], id_bits: int) -> tuple[str, str]:
"""The channel letter and the sensor identity, as the family encodes it."""
channel = CHANNELS[(data[0] >> 6) & 0x03]
mask = (1 << (id_bits - 8)) - 1
return channel, f"{((data[0] & mask) << 8) | data[1]:04X}"
def decode_tower(data: list[int]) -> Reading | None:
"""The 592TXR / 06002RM outdoor sensor: temperature and humidity.
Fourteen bits of identity with the channel above them, a status byte
whose bit 6 is set while the battery is good, humidity in seven bits, and
temperature in eleven -- tenths of a degree Celsius offset by a hundred,
so that the coldest thing it can report is still a positive number.
This one has been held against messages off real sensors rather than only
against messages built here; see the tests.
"""
if len(data) != 7 or not _txr_framed(data):
return None
if (data[2] & 0x3F) != 0x04:
return _unknown("tower", data, 14)
channel, sensor = _txr_identity(data, 14)
humidity = data[3] & 0x7F
raw = ((data[4] & 0x0F) << 7) | (data[5] & 0x7F)
celsius = raw / 10.0 - 100.0
if not _sane_temperature(celsius) or not _sane_humidity(humidity):
return None
return Reading(
model="Tower 592TXR", family="tower", sensor=sensor, channel=channel,
battery_low=not data[2] & 0x40, message=data[2] & 0x3F,
measures=(Measure("temperature", round(celsius, 1), "C"),
Measure("humidity", float(humidity), "%")),
checks=("checksum-8", "parity"))
def decode_5n1(data: list[int]) -> Reading | None:
"""The 5-in-1, which says half of what it knows at a time.
It alternates two messages, because there is more to say than fits in one
of them: a wind message with the direction and the rain counter, and a
weather message with the temperature and the humidity. Both carry the
wind speed, which is the one number worth having on every transmission.
"""
if len(data) != 8 or not _txr_framed(data):
return None
kind = data[2] & 0x3F
channel, sensor = _txr_identity(data, 12)
common = dict(model="5-in-1 06014RM", family="5n1", sensor=sensor,
channel=channel, battery_low=not data[2] & 0x40,
message=kind, checks=("checksum-8", "parity"))
# The same eight bits in both messages: five from one byte and three
# from the next. Zero means stopped; anything else is a count the
# published description turns into kilometres an hour with a slope and
# an offset, and the offset is why a turning cup never reads zero.
raw_wind = ((data[3] & 0x1F) << 3) | ((data[4] & 0x70) >> 4)
wind = 0.0 if raw_wind == 0 else raw_wind * 0.8278 + 1.0
if not WIND_RANGE[0] <= wind <= WIND_RANGE[1]:
return None
if kind == 0x31: # wind, where from, and rain
degrees = WIND_POINTS[data[4] & 0x0F]
counter = ((data[5] & 0x7F) << 7) | (data[6] & 0x7F)
return Reading(
measures=(Measure("wind", round(wind, 1), "km/h"),
Measure("wind from", degrees, "deg"),
Measure("rain", round(counter * RAIN_PER_COUNT_MM, 2),
"mm", raw=float(counter))),
**common)
if kind == 0x38: # wind, temperature and humidity
raw = ((data[4] & 0x0F) << 7) | (data[5] & 0x7F)
celsius = _fahrenheit((raw - 400) / 10.0)
humidity = data[6] & 0x7F
if not _sane_temperature(celsius) or not _sane_humidity(humidity):
return None
return Reading(
measures=(Measure("temperature", round(celsius, 1), "C"),
Measure("humidity", float(humidity), "%"),
Measure("wind", round(wind, 1), "km/h")),
**common)
return _unknown("5n1", data, 12)
def decode_lightning(data: list[int]) -> Reading | None:
"""The 6045M, which is a tower sensor that also counts lightning.
The strike counter is cumulative and wraps at 127, and the distance is
the storm's, in miles, as the detector estimates it -- zero to thirty-one,
where thirty-one means "further away than this can tell". There is also
a bit which is set when the detector believes it is being interfered with,
and it is worth showing, because a strike count that climbs while that
bit is set is not lightning.
The message type is insisted on, and that is the whole reason this can be
run at all. The Atlas is a nine-byte sensor too and lays its payload out
differently, and nothing but the type field tells them apart: decoded as
this, an Atlas would give a temperature read off the wrong bits, which
would pass the checksum, pass the parity, and land inside a plausible
range often enough to be believed. Wrong weather under somebody's sensor
name is a worse answer than none, so anything else nine bytes long comes
back as a sensor of a model this cannot read.
"""
if len(data) != 9 or not _txr_framed(data):
return None
kind = data[2] & 0x3F
if kind != LIGHTNING_MESSAGE:
return _unknown("6045", data, 12)
channel, sensor = _txr_identity(data, 12)
humidity = data[3] & 0x7F
raw = ((data[4] & 0x1F) << 7) | (data[5] & 0x7F)
celsius = _fahrenheit(raw / 10.0 - 100.0)
if not _sane_temperature(celsius) or not _sane_humidity(humidity):
return _unknown("6045", data, 12)
strikes = data[6] & 0x7F
miles = data[7] & 0x1F
measures = [Measure("temperature", round(celsius, 1), "C"),
Measure("humidity", float(humidity), "%"),
Measure("strikes", float(strikes), "")]
if miles < 0x1F:
measures.append(Measure("storm", round(miles * 1.609344, 1), "km",
raw=float(miles)))
flags = ["checksum-8", "parity"]
if data[7] & 0x20:
flags.append("interference")
return Reading(
model="Lightning 6045M", family="6045", sensor=sensor, channel=channel,
battery_low=not data[2] & 0x40, message=kind,
measures=tuple(measures), checks=tuple(flags))
def _unknown(family: str, data: list[int], id_bits: int) -> Reading:
"""A message that framed correctly and is not one this can read.
Reported rather than dropped, because knowing that there is a sensor of
some kind on the fence, with an identity that stays the same, is worth
something on its own -- it can be given a name and watched -- and because
silently discarding a well-formed message is how a decoder comes to look
like a receiver problem.
"""
channel, sensor = _txr_identity(data, id_bits)
return Reading(model=f"AcuRite {len(data)}-byte sensor",
family=family, sensor=sensor, channel=channel,
battery_low=not data[2] & 0x40, message=data[2] & 0x3F,
measures=(), checks=("checksum-8", "parity"))
# ---------------------------------------------------------------------------
# The two older ones
# ---------------------------------------------------------------------------
def _signed12(high: int, low: int) -> int:
"""Twelve bits of two's-complement temperature, in tenths of a degree."""
raw = ((high & 0x0F) << 8) | low
return raw - 0x1000 if raw & 0x800 else raw
def decode_609(data: list[int]) -> Reading | None:
"""The 609TXC: one byte of identity, temperature, humidity, a sum.
The identity is only eight bits and is redrawn at random every time the
battery is changed, so two of these in one garden will collide about once
in every two hundred and fifty battery changes. Nothing can be done
about that from here; it is worth knowing when a sensor appears to have
changed its mind about the weather.
"""
if len(data) != 5 or (sum(data[:4]) & 0xFF) != data[4] or not any(data):
return None
celsius = _signed12(data[1], data[2]) / 10.0
humidity = data[3]
if not _sane_temperature(celsius) or not _sane_humidity(humidity):
return None
return Reading(
model="609TXC", family="609", sensor=f"{data[0]:02X}",
battery_low=not data[1] & 0x80,
measures=(Measure("temperature", round(celsius, 1), "C"),
Measure("humidity", float(humidity), "%")),
checks=("checksum-8", "twice"))
def decode_606(data: list[int]) -> Reading | None:
"""The 606TX: identity, temperature, a CRC-8, and nothing else at all."""
if len(data) != 4 or crc8(data[:3]) != data[3] or not any(data):
return None
celsius = _signed12(data[1], data[2]) / 10.0
if not _sane_temperature(celsius):
return None
return Reading(
model="606TX", family="606", sensor=f"{data[0]:02X}",
battery_low=not data[1] & 0x80,
measures=(Measure("temperature", round(celsius, 1), "C"),),
checks=("crc-8", "twice"))
def _sane_temperature(celsius: float) -> bool:
return TEMPERATURE_RANGE[0] <= celsius <= TEMPERATURE_RANGE[1]
def _sane_humidity(humidity: float) -> bool:
return HUMIDITY_RANGE[0] <= humidity <= HUMIDITY_RANGE[1]
_DECODERS = ((7, decode_tower), (8, decode_5n1), (9, decode_lightning),
(5, decode_609), (4, decode_606))
# ---------------------------------------------------------------------------
# From bits to readings
# ---------------------------------------------------------------------------
# Which end of a byte goes down the air first. Most significant bit first is
# what the descriptions of these formats assume, and it is what nearly
# everything does, but "nearly" is the word that has cost this section three
# rounds already -- so both are tried and the checksums say which.
#
# There is a signature worth knowing for this one. Reversing the bits of a
# byte does not change how many of them are set, so parity survives it and a
# sum does not: a message read from the wrong end shows its parity holding
# and its checksum failing, every time, on every copy.
BIT_ORDERS = (False, True)
_REVERSED = [int(f"{byte:08b}"[::-1], 2) for byte in range(256)]
def _bytes_at(bits: str, at: int, count: int,
reflect: bool = False) -> list[int]:
out = [int(bits[at + i * 8:at + (i + 1) * 8], 2) for i in range(count)]
return [_REVERSED[byte] for byte in out] if reflect else out
def candidates(bits: str) -> list[Reading]:
"""Every message that can be read out of a run of bits, at any offset.
Every offset, because what arrives here has a sync pattern of some length
in front of it and the message does not begin on a byte boundary of the
recovered bits. Every length, because a burst does not say which model
sent it. The checks are what make that affordable: a message which
satisfies a sum, a parity set and a plausibility range at a bit offset it
does not belong at is rare enough to be worth the search.
"""
found: list = []
for count, decoder in _DECODERS:
# A message has to sit where a message sits: a few bits of sync in
# front of it, and the end of the burst immediately behind it. This
# is the check that stops a short model being read out of a long one.
#
# It is needed because neither of the other two defences reaches the
# case. Corroboration does not: the copies of a message are
# identical, so a coincidence in one copy is a coincidence in all
# three, and asking for it three times over asks for nothing. Parity
# does not either, and this is the part that is easy to get wrong --
# a seven-byte window taken from the front of a real eight-byte
# message is made of that message's own payload bytes, whose parity
# is already correct, so the four parity bits contribute nothing at
# all and what is left is one byte of sum. One in two hundred and
# fifty-six is not a rate at which a display can be trusted.
#
# Where it sits, though, gives that window away every time: it ends
# a whole byte before the burst does.
need = count * 8
for at in range(0, len(bits) - need + 1):
if at > LEAD_BITS or len(bits) - at - need > TAIL_BITS:
continue
if not _fits(at, need, len(bits)):
continue
for reflect in BIT_ORDERS:
try:
data = _bytes_at(bits, at, count, reflect)
except ValueError:
break
reading = decoder(data)
if reading is None:
continue
if not reading.measures and not _fits(at, need, len(bits),
unread=True):
continue
reading.bits = bits[at:at + need]
reading.offset = at
found.append((at, need, reading))
break # one order or the other, not both
return _one_per_window(found, len(bits))
def _fits(at: int, need: int, length: int, unread: bool = False) -> bool:
"""Whether a message of this length can sit at this offset in a burst.
A message begins after the sync and runs to the end of the burst, so a
handful of bits in front of it and almost none behind. One that cannot
be read is held to the tighter figures: it has no type field that means
anything here and no range that can be checked, so where it sits is the
only evidence there is that it is a message at all.
"""
lead = UNREAD_LEAD_BITS if unread else LEAD_BITS
tail = UNREAD_TAIL_BITS if unread else TAIL_BITS
return at <= lead and length - at - need <= tail
def _one_per_window(found: list, length: int) -> list[Reading]:
"""Of two messages sharing bits, keep one: they cannot both be real.
A message forty bits long, found in a burst forty-one bits long, can be
read at two offsets, and one byte of checksum is satisfied by the wrong
one of them about once in every two hundred and fifty-six bursts. Both
survive corroboration, because both are in all three copies. What tells
them apart is that a real message ends where the burst does, and the
other one leaves a bit hanging off the end.
Longer messages and better-understood ones are preferred first, so that
a burst which really does hold an eight-byte message is never given up
for a five-byte window inside it.
"""
found.sort(key=lambda item: (len(item[2].measures), item[1],
-(length - item[0] - item[1]), -item[0]),
reverse=True)
kept: list[Reading] = []
taken: list[tuple[int, int]] = []
for at, need, reading in found:
if any(at < end and start < at + need for start, end in taken):
continue
taken.append((at, at + need))
kept.append(reading)
return kept
def decode(bits: str, confirm: bool = True) -> Reading | None:
"""The one message a run of bits holds, or None.
Where a model's checks are thin -- the two older ones carry a single
byte between them -- the same message has to turn up more than once
before it is believed, and one run of bits is one copy. So by default
this reads the three well-checked models and not the two thin ones; the
thin ones are reached through :func:`readings_from`, which sees a whole
block and so can see the copies.
``confirm=False`` reads whatever framed correctly, for a caller with its
own reason to believe a single message: a stored log, or a test holding
the encoder that made it.
"""
return _best(candidates(bits), confirm)
def _best(found: list[Reading], confirm: bool = True) -> Reading | None:
"""The most-corroborated message in a pile of candidates.
Ordered by how many copies arrived, then by how much check the model
carries, then by how long the message is -- a longer message that framed
correctly is better evidence than a shorter one that did, and the
longest ones are also the ones with the most to say.
"""
if not found:
return None
seen: dict[tuple, list[Reading]] = {}
for reading in found:
seen.setdefault((reading.family, reading.bits), []).append(reading)
good = [(copies, group[0]) for (family, _bits), group in seen.items()
if (copies := len(group)) >= (REPEATS_NEEDED[family] if confirm
else 1)]
if not good:
return None
good.sort(key=lambda pair: (pair[0], len(pair[1].bits),
len(pair[1].measures)), reverse=True)
return good[0][1]
def decode_burst(burst: "Burst") -> Reading | None:
"""One burst of pulses, read as both codings, decoded as whatever it is.
These sensors do not agree on how a bit is drawn. The newer ones vary
the length of the pulse and keep the gap between pulses roughly even;
the two older ones keep the pulse even and vary the gap. Rather than
decide which from the timings -- which is guessing, and wrong on a weak
burst where the edges have moved -- both readings of the same pulses are
tried, and the checksums say which one it was.
"""
return _best(_from_burst(burst))
def _from_burst(burst: "Burst") -> list[Reading]:
"""The messages in one burst, from the first reading of it that yields any.
The readings are tried in order of likelihood and the search stops at the
first one that produces anything, rather than pooling all of them. That
matters more than it looks.
Half a dozen readings of a burst, each tried at both byte orders, is
sixteen times as many chances for a coincidence to satisfy a checksum as
one reading was -- and the checks here are twelve to fourteen bits, which
is a rate that is comfortable once and uncomfortable sixteen times over.
Sensors that are not there began appearing in the invented garden the
moment the alternatives went in.
Stopping early costs nothing. A burst that reads correctly under the
ordinary reading never reaches the alternatives; a burst that does not
reach them exactly as it did before. The fallbacks are there for the
sensor whose bits are drawn some other way, and that sensor is not in
competition with anything.
"""
for bits in slicings(burst):
found = candidates(bits)
if found:
return found
return []
# ---------------------------------------------------------------------------
# From the air to bits: an on-off-keyed envelope, sliced
# ---------------------------------------------------------------------------
@dataclass
class Burst:
"""One run of pulses with no long silence in it.
Marks and spaces alternate and are in microseconds, starting with a mark;
there is always one more mark than there are spaces, because the silence
that ended the burst is not part of it.
"""
at: float = 0.0 # seconds from the start of the block
marks: tuple = ()
spaces: tuple = ()
level: float = 0.0 # how far above the noise floor, as a ratio
@property
def pulses(self) -> int:
return len(self.marks)
@property
def length_us(self) -> float:
return float(sum(self.marks) + sum(self.spaces))
def baseband(iq: np.ndarray, sample_rate: float, offset: float = 0.0,
want_rate: float = 250_000.0) -> tuple[np.ndarray, float]:
"""The envelope of the sensor's channel, at a rate worth slicing.
Three things happen here, in this order and for this reason.
The receiver is tuned a little to one side of 433.92 MHz, because every
RTL-SDR puts a spike of its own at whatever it is tuned to and a spike
sitting on top of an on-off-keyed signal is the one thing that stops it
being on and off. So the first step is to shift the sensor back to the
middle, which puts the spike out at the edge instead.
The second is a plain running average of the complex samples, long enough
that its first null lands on the spike. It is a crude filter and a
deliberately crude one: it is four instructions wide, it runs on a whole
second of samples without noticing, and what it has to reject is one
tone at a frequency this end chose.
Only then is the magnitude taken. Filtering before detection rather than
after is what keeps the neighbours -- a doorbell, a tyre sensor, a car
key -- from adding themselves to the envelope of the sensor.
"""
if iq.size == 0:
return np.zeros(0, dtype=np.float32), float(sample_rate)
signal = np.asarray(iq)
if offset:
turn = np.exp(-2j * math.pi * offset
* np.arange(signal.size, dtype=np.float64) / sample_rate)
signal = signal * turn.astype(np.complex64)
decimate = max(1, int(sample_rate // max(1.0, want_rate)))
if decimate > 1:
usable = (signal.size // decimate) * decimate
if usable == 0:
return np.zeros(0, dtype=np.float32), float(sample_rate)
signal = signal[:usable].reshape(-1, decimate).mean(axis=1)
return np.abs(signal).astype(np.float32), float(sample_rate) / decimate
def bursts(envelope: np.ndarray, rate: float, gap_us: float = 3_000.0,
min_pulses: int = 8, min_run_us: float = 70.0,
min_level: float = 1.8) -> list[Burst]:
"""Split an envelope into the runs of pulses worth trying to read.
This is done in two passes, and the reason is a garden with more than one
sensor in it.
The first pass only asks where anything is happening at all, and asks it
against the noise rather than against the loudest thing in the block.
That distinction is the whole point. A single threshold set halfway
between the noise floor and the peak of a whole second sounds reasonable
and is not: a sensor on the windowsill and a sensor at the end of the
garden differ by forty decibels, so a threshold set halfway to the near
one sits far above everything the far one ever does, and the far one
disappears -- not weakly, but completely, and only while the near one is
transmitting. A block with one loud sensor in it would yield exactly one
sensor however many were out there.
So the gate here is the noise floor plus a few times its own scatter,
which is a level the quietest readable burst still clears and noise does
not. Whatever clears it is grouped into regions, and the second pass
then re-thresholds each region against its own high and low. Every
sensor is sliced at its own amplitude, and a burst forty decibels down on
its neighbour is read exactly as well as the neighbour is.
"""
if envelope.size < 4:
return []
smooth = _smoothed(envelope, rate)
# The noise, read off the bottom of the block rather than the middle of
# it, and for the same reason the gate is: the middle of a block that is
# largely signal is signal. Taken from the median, a burst occupying
# most of what it was handed looks no louder than the thing it is
# measured against, and the whole block is thrown away as empty -- which
# is what happened to any capture short enough to be mostly burst.
quiet = float(np.percentile(smooth, 20))
peak = float(np.percentile(smooth, 99.99))
if peak <= quiet * min_level or peak <= 0.0:
return [] # nothing above the noise worth slicing
hot = smooth > _noise_gate(smooth, quiet, peak)
if not hot.any():
return []
per_us = rate / 1e6
out: list[Burst] = []
# Grouped tightly, then joined back up. Tightly, because the region is
# what fixes the threshold, and a region holding two sensors of unequal
# strength fixes it on the louder -- which is the fault this whole
# arrangement exists to avoid, and it comes straight back if regions are
# allowed to be generous.
#
# But a tight region cannot hold a whole message from every model: the
# gaps inside one range from two hundred microseconds on the newer
# sensors to four thousand on the oldest, and a grouping tight enough to
# keep two sensors apart tears the oldest into a bit at a time.
#
# So the joining is done afterwards, on the sliced pulses rather than on
# the samples, where each piece has already been measured at its own
# amplitude and the gaps are known. Two pieces belong to one message
# when the silence between them looks like the silences inside them.
for lo, hi in _regions(hot, int(round(gap_us * per_us))):
out += _slice(smooth, lo, hi, rate, gap_us, min_run_us, quiet)
out.sort(key=lambda b: b.at)
return [b for b in _divide(_join(out)) if b.pulses >= min_pulses]
def _join(found: list[Burst]) -> list[Burst]:
"""Put back together the pieces of one message, and no more than that.
The test is whether the silence between two pieces looks like the
silences within them. Inside a message every gap is much of a size;
between one copy of a message and the next it is several times that. So
a separation of the same order as the gaps either side of it is part of
the message, and one much larger is the wait for the next copy.
"""
# What a gap inside a message looks like across the whole block, for
# pieces too small to say. The first pulses of a message often come off
# as a piece of one pulse and no gaps at all, and judged on their own
# contents there is nothing to judge -- so they were left stranded, and a
# message short of its first two bits is a message short of its checksum.
everything = [gap for burst in found for gap in burst.spaces]
fallback = float(np.median(everything)) if everything else 0.0
out: list[Burst] = []
for burst in found:
if out:
last = out[-1]
apart = (burst.at - (last.at + last.length_us / 1e6)) * 1e6
gaps = list(last.spaces) + list(burst.spaces)
typical = float(np.median(gaps)) if gaps else fallback
if 0.0 <= apart <= max(2.5 * typical, 700.0):
out[-1] = Burst(
at=last.at,
marks=last.marks + burst.marks,
spaces=last.spaces + (apart,) + burst.spaces,
level=max(last.level, burst.level))
continue
out.append(burst)
return out
def _smoothed(envelope: np.ndarray, rate: float) -> np.ndarray:
"""A running mean over a fraction of the shortest pulse these send.
Enough that a sample of noise on the wrong side of a threshold cannot
become a pulse, and little enough that the shortest real pulse -- about
two hundred microseconds -- survives it with its edges where they were.
"""
span = max(1, int(round(rate * 30e-6)))
if span <= 1:
return envelope
pad = np.concatenate(([0.0], np.cumsum(envelope, dtype=np.float64)))
return ((pad[span:] - pad[:-span]) / span).astype(np.float32)
def _noise_gate(smooth: np.ndarray, quiet: float, peak: float) -> float:
"""The level below which nothing is worth looking at.
This only has to find where something happened; how loud it was is
settled afterwards, region by region. So it wants an estimate of the
noise and nothing else, and in particular it must not move when
something loud happens -- that is exactly the case where a sensor at the
end of the garden is about to be lost.
Everything here is read off the bottom of the distribution rather than
the middle of it. A second of band holds a few hundred milliseconds of
sensor at the very most, so the lowest fifth of it is noise however busy
the rest was. The middle is not safe in the same way: the median and the
deviation about it both climb once a fair fraction of the block is
signal, and a gate built on them can end up above the peak and pass
nothing at all.
The other two terms are floors under the gate, for a signal so clean that
the noise has almost no scatter to measure -- a synthetic envelope, or a
receiver with nothing connected -- where the first term alone would be
zero and everything would be above it.
Note what none of the three terms is: a fraction of the way to the peak.
That is the obvious way to write this and it is the bug it replaced. Any
gate proportional to the loudest thing in the block rises when a sensor
close to the aerial transmits, and rises past every quieter sensor out
there -- so the far ones vanish, and vanish only while the near one is
talking, which is as confusing a symptom as radio produces.
"""
scatter = float(np.percentile(smooth, 40) - np.percentile(smooth, 10))
gate = quiet + max(6.0 * scatter, 0.5 * quiet, 0.001 * (peak - quiet))
# And never above the loudest thing in the block, whatever the arithmetic
# above comes to. A gate over the peak is the one setting that cannot be
# right: it reports silence on a second that plainly had something in it.
#
# It binds when the carrier is on for most of what was handed over -- a
# sensor with short gaps, or a capture trimmed close around one -- where
# the percentiles the scatter is taken from straddle the edge between off
# and on, and six times the resulting figure is a gate above everything
# in the block. Reading those percentiles lower down avoids that case
# instead of catching it, and was tried; but there turns out to be no
# signal it rescues that this does not, because a block cannot be mostly
# one sensor and also hold a much quieter one, so the simpler of the two
# is what is here.
return min(gate, quiet + 0.8 * (peak - quiet))
def _regions(hot: np.ndarray, gap_samples: int) -> list[tuple[int, int]]:
"""Spans of activity, with short silences inside them kept.
The silences between the pulses of one message are part of the message,
so they must not end a region; only a silence longer than any gap within
a message does that.
"""
lengths, values, starts = _runs(hot)
spans: list[list[int]] = []
for length, live, start in zip(lengths, values, starts):
if not live:
continue
if spans and start - spans[-1][1] <= gap_samples:
spans[-1][1] = start + length
else:
spans.append([start, start + length])
return [(lo, hi) for lo, hi in spans]
def _slice(smooth: np.ndarray, lo: int, hi: int, rate: float, gap_us: float,
min_run_us: float, block_floor: float) -> list[Burst]:
"""Cut one region into marks and spaces, at its own amplitude.
The threshold is halfway between what this burst does when it is on and
what it does when it is off, both taken from the region itself. A little
of the silence either side is included so that there is an "off" to
measure even when the region is nearly all pulses.
"""
per_us = rate / 1e6
margin = int(round(400 * per_us))
lo = max(0, lo - margin)
hi = min(smooth.size, hi + margin)
part = smooth[lo:hi]
if part.size < 4:
return []
off = float(np.percentile(part, 5))
on_level = float(np.percentile(part, 98))
if on_level <= off:
return []
on = part > (off + on_level) / 2.0
lengths, values, starts = _runs(on)
lengths, values, starts = _merge_short(lengths, values, starts,
min_run_us * per_us)
out: list[Burst] = []
marks: list[float] = []
spaces: list[float] = []
began = 0
ends = _copy_gap(lengths, values, per_us, gap_us)
# A silence only becomes a bit's gap once another pulse follows it. The
# silence at the end of a burst is not part of the last bit -- it is the
# wait until the next copy, and it is as long as that wait happens to be
# -- so measuring the last bit against it reads the last bit wrongly,
# which is the last bit of the checksum, which is the whole message.
# Held back rather than appended, and dropped if nothing follows.
waiting: float | None = None
for length, high, start in zip(lengths, values, starts):
micro = length / per_us
if high:
if not marks:
began = start
elif waiting is not None:
spaces.append(waiting)
waiting = None
marks.append(micro)
elif marks:
if micro > ends:
out.append(_burst_of(marks, spaces, lo + began, rate,
block_floor, on_level))
marks, spaces, waiting = [], [], None
else:
waiting = micro
if marks:
out.append(_burst_of(marks, spaces, lo + began, rate, block_floor,
on_level))
return out
def _copy_gap(lengths, values, per_us: float, ceiling: float) -> float:
"""The silence that means "that was the whole message", in microseconds.
Measured from the message rather than fixed, because the gaps inside one
differ by model: two hundred microseconds on the newer sensors, four
thousand on the oldest. A fixed figure has to be larger than the largest
gap inside any message and smaller than the smallest gap between two
copies of one, and there is no such figure -- pick it high and the three
copies of a message merge into one burst that matches nothing, pick it
low and a single message is torn into a bit at a time.
So it is taken from the burst. From the widest gap in it rather than
the middle one, which is the part that is easy to get wrong: where a
message draws a one as a gap three times the length of a zero, and most
of the bits are zeroes, the middle gap is the short one and twice it
still falls inside the message -- so every one-bit ends a burst and the
message comes apart into pieces of two or three pulses. The widest gap
inside a message is a fact about the message; the middle one is a fact
about the data it happens to be carrying that day.
"""
gaps = [length / per_us for length, live in zip(lengths, values)
if not live]
if len(gaps) < 4:
return ceiling
widest = float(np.percentile(gaps, 90))
return min(ceiling, max(2.0 * widest, 700.0))
# The longest message here is nine bytes, which with the four sync pulses in
# front of it is seventy-six. Anything appreciably longer than that is more
# than one message however it came to be in one piece.
MAX_PULSES = 88
def _divide(found: list[Burst]) -> list[Burst]:
"""Cut apart anything too long to be a single message.
Joining pieces back together goes by whether the silence between them
looks like the silences within them, and on the oldest sensor those two
are only a factor of two apart -- four thousand microseconds inside a
message and eight thousand between copies of it -- which is too close to
call from a ratio. So the join is allowed to be greedy and this undoes
it where the result is plainly too long to be one message: the largest
gaps in such a burst are the boundaries between the copies, because that
is what they physically are.
Kept separate from the joining rather than folded into it because the
two answer different questions. The join asks whether two pieces belong
together and can be wrong in one direction only; this asks how many
messages are in front of it, which is a question that can be answered by
counting.
"""
out: list[Burst] = []
for burst in found:
if burst.pulses <= MAX_PULSES or len(burst.spaces) < 8:
out.append(burst)
continue
cut = 1.5 * float(np.percentile(burst.spaces, 90))
pieces, marks, spaces = [], [], []
at = burst.at
began = at
for i, mark in enumerate(burst.marks):
marks.append(mark)
at += mark / 1e6
gap = burst.spaces[i] if i < len(burst.spaces) else None
if gap is None:
continue
if gap > cut:
pieces.append(Burst(at=began, marks=tuple(marks),
spaces=tuple(spaces), level=burst.level))
marks, spaces = [], []
began = at + gap / 1e6
else:
spaces.append(gap)
at += gap / 1e6
if marks:
pieces.append(Burst(at=began, marks=tuple(marks),
spaces=tuple(spaces), level=burst.level))
out += pieces if len(pieces) > 1 else [burst]
return out
def _burst_of(marks, spaces, began, rate, floor, peak) -> Burst:
"""One burst, from the marks and spaces it was cut into.
There is always one more mark than there are spaces: the silence that
ended the burst belongs to whatever comes after it, not to this.
Everything is cast to an ordinary float on the way out. These numbers
came from numpy and end up as a timestamp in a log and in a file of
names, and neither JSON nor YAML will write a numpy float -- which is the
sort of thing that only shows up when a sensor is finally named at two in
the morning.
"""
marks = [float(m) for m in marks][:len(spaces) + 1]
return Burst(at=float(began) / float(rate), marks=tuple(marks),
spaces=tuple(float(s) for s in spaces),
level=float(peak / floor) if floor > 0 else float("inf"))
def _runs(on: np.ndarray):
"""Every run of equal values: how long, which value, where it started."""
change = np.flatnonzero(np.diff(on.view(np.int8)))
edges = np.concatenate(([0], change + 1, [on.size]))
return np.diff(edges), on[edges[:-1]], edges[:-1]
def _merge_short(lengths, values, starts, floor_samples: float):
"""Fold runs too short to be a pulse into whatever is around them.
A single sample the wrong side of the threshold in the middle of a mark
would otherwise split one four-hundred-microsecond pulse into three, and
everything after that is arithmetic on nonsense. Smoothing removes most
of it; this removes the rest.
"""
if lengths.size == 0:
return lengths, values, starts
keep = lengths >= max(1.0, floor_samples)
if keep.all():
return lengths, values, starts
if not keep.any():
return lengths[:0], values[:0], starts[:0]
out_len, out_val, out_start = [], [], []
for length, value, start in zip(lengths[keep], values[keep], starts[keep]):
if out_val and out_val[-1] == value:
# Two runs of the same value now touching: the short run between
# them was noise, and its samples belong to the pulse it split.
out_len[-1] = start + length - out_start[-1]
else:
out_len.append(int(length))
out_val.append(bool(value))
out_start.append(int(start))
return (np.array(out_len), np.array(out_val, dtype=bool),
np.array(out_start))
def bits_pwm(burst: Burst) -> str:
"""Read a burst as pulse width: a long mark is a one, a short one a zero.
Nothing here is measured against a clock. The bit is decided by
comparing each mark with the gap that follows it, which is a comparison
between two numbers that arrived a few hundred microseconds apart and so
cannot have drifted with respect to each other. A transmitter running
ten per cent fast is read correctly and never noticed.
The last pulse of a burst is the exception, and has to be handled or the
final bit of the checksum is lost and with it the message: the gap after
it is the silence before the next copy rather than part of the bit. What
stands in for that gap is the bit period -- a mark and its gap together,
which in this coding is the same length whichever bit it is -- taken as
the median over the burst. A pulse longer than half a bit period is a
one, which is what the coding means.
"""
if not burst.marks:
return ""
spaces = list(burst.spaces)
if len(burst.marks) > len(spaces):
if not spaces:
return ""
period = float(np.median([mark + space for mark, space
in zip(burst.marks, spaces)]))
spaces.append(period / 2.0)
return "".join("1" if mark > space else "0"
for mark, space in zip(burst.marks, spaces))
def bits_ppm(burst: Burst) -> str:
"""Read a burst as pulse position: a long gap is a one, a short one zero.
Here there is nothing to compare a gap against except the other gaps, so
the threshold is halfway between the shortest and the longest of them.
A burst whose gaps are all much the same length is not this coding, and
comes back as an empty string rather than as a row of zeroes that
something downstream might believe.
"""
if len(burst.spaces) < 2:
return ""
low, high = min(burst.spaces), max(burst.spaces)
if high < low * 1.5:
return ""
middle = (low + high) / 2.0
return "".join("1" if space > middle else "0" for space in burst.spaces)
def _levels(values, most: int = 3) -> list[float]:
"""Thresholds that might separate short from long, best division first.
The lengths in a burst fall into clusters -- two of them for the bits,
and often a third for the sync, which is longer than either. Rather than
decide which cluster boundary is the one that means something, this
returns a few of the widest gaps in the sorted lengths and lets the
checksums say. Guessing here is what a fixed threshold does, and a fixed
threshold is wrong the moment a transmitter warms up.
"""
ordered = sorted(float(v) for v in values)
if len(ordered) < 4:
return []
span = ordered[-1] - ordered[0]
if span <= 0:
return []
steps = sorted(((ordered[i + 1] - ordered[i], i)
for i in range(len(ordered) - 1)), reverse=True)
out = []
for step, i in steps[:most]:
if step < 0.08 * span:
break # not a division, just jitter
out.append((ordered[i] + ordered[i + 1]) / 2.0)
return sorted(out)
_FLIP = str.maketrans("01", "10")
def _by_width(values, level: float) -> str:
return "".join("1" if value > level else "0" for value in values)
def slicings(burst: Burst) -> list[str]:
"""Every way this burst might be read, for the checksums to choose from.
There are two things not to assume about a burst, and both of them are
assumptions that hold on the sensor in front of you and fail on the next
one.
The first is which of the pulse and the gap carries the bit. The newer
models vary the pulse; the two older ones vary the gap.
The second is subtler and is the one that had this reading nothing at
all. Where the pulse carries the bit, the gap may be the complement of
it -- so that every bit takes the same time, and a pulse can be judged
against the gap that follows it -- or the gap may simply be a fixed
spacer. Judged against a fixed spacer of two hundred microseconds, a
short pulse of two hundred and twenty is longer than its gap and reads as
a one, which is the wrong bit, and every message fails its checksum with
nothing to say why.
So nothing is assumed. The pulse against its own gap, the pulse against
each threshold the pulse lengths themselves suggest, and the same for the
gaps: half a dozen readings of the same burst, of which at most one will
satisfy a checksum. It costs a few microseconds per burst and it is the
difference between working on the sensors this was written against and
working on the ones in somebody's garden.
"""
out = [bits_pwm(burst), bits_ppm(burst)]
for values in (burst.marks, burst.spaces):
for level in _levels(values):
# And the other way up. Which of long and short means one is a
# third thing not to assume, and complementing a bit string is
# free.
reading = _by_width(values, level)
out += [reading, reading.translate(_FLIP)]
seen: set = set()
keep = []
for bits in out:
if bits and bits not in seen:
seen.add(bits)
keep.append(bits)
return keep
def readings_from(iq: np.ndarray, sample_rate: float, offset: float = 0.0,
when: float = 0.0) -> list[Reading]:
"""Every distinct sensor message in one block of samples, in order.
``when`` is the clock time the block began, so that each reading carries
the real moment it arrived rather than its offset in a buffer. A log of
weather that started again from zero would be unusable, and a sensor
reports every sixteen seconds, so the difference matters.
The three copies of a message are looked for across the whole block
rather than inside one burst, because they are not one burst: these
sensors leave ten milliseconds of silence between the copies, which is
three times what ends a burst. Corroboration therefore has to happen
here, where the whole second is in view, and a message that arrived three
times comes back as one reading which knows it arrived three times.
"""
envelope, rate = baseband(iq, sample_rate, offset)
found = []
for burst in bursts(envelope, rate):
# One burst is one copy of one message, however many ways of reading
# it happened to produce it. Counted per reading rather than per
# slicing, or two readings of the same burst would corroborate each
# other -- and corroboration is the only thing standing between the
# two thinly-checked models and a display full of sensors that are
# not there.
here: dict = {}
for reading in _from_burst(burst):
here.setdefault((reading.family, reading.bits), reading)
for reading in here.values():
reading.at = when + burst.at
found.append(reading)
return confirmed(found)
def confirmed(found: list[Reading]) -> list[Reading]:
"""One reading per distinct message, once enough copies have arrived.
Identical bits are one message heard more than once; different bits are
different messages, even from the same sensor, which is how the 5-in-1's
two halves both survive a block that holds them both.
"""
seen: dict[tuple, list[Reading]] = {}
for reading in found:
seen.setdefault((reading.family, reading.bits), []).append(reading)
out = []
for (family, _bits), group in seen.items():
if len(group) < REPEATS_NEEDED[family]:
continue
first = min(group, key=lambda r: r.at)
first.copies = len(group)
out.append(first)
return sorted(out, key=lambda r: (r.at, r.family, r.sensor))
# ---------------------------------------------------------------------------
# Looking at what actually arrived
# ---------------------------------------------------------------------------
def timings(values, tolerance: float = 0.25) -> list[tuple[float, int]]:
"""The distinct lengths in a burst, each with how many there were.
A message drawn as pulses has two or three lengths in it and nothing in
between, so this is the shape of the protocol written out: (220, 28)
(402, 28) (598, 4) is a burst of fifty-six bits and four sync pulses,
and anyone who knows the sensor can see that at a glance. It is the
single most useful thing to print when nothing is decoding, because it
says whether the trouble is the radio or the arithmetic.
"""
ordered = sorted(float(v) for v in values)
groups: list[list[float]] = []
for value in ordered:
if groups and value <= groups[-1][-1] * (1.0 + tolerance):
groups[-1].append(value)
else:
groups.append([value])
return [(sum(g) / len(g), len(g)) for g in groups]
@dataclass(frozen=True)
class NearMiss:
"""One way of framing a burst, and which of its checks held.
"Nothing framed" is not an answer, it is the absence of one. A message
that satisfies its checksum and fails its parity is a different fault
from one that fails both, and both are different from a burst that never
lined up on a byte boundary at all -- and the three want three different
fixes. This is what turns the first into the second.
"""
family: str = ""
length: int = 0 # in bytes
offset: int = 0 # where in the burst it starts, in bits
data: tuple = ()
checks: tuple = () # (name, passed) in the order they are run
read: bool = False # whether the decoder accepted it outright
reflect: bool = False # whether the bytes were read the other way
@property
def score(self) -> int:
return sum(1 for _name, passed in self.checks if passed)
@property
def hex(self) -> str:
return " ".join(f"{byte:02X}" for byte in self.data)
def describe(self) -> str:
marks = " ".join(("\u2713" if passed else "\u2717") + name
for name, passed in self.checks)
order = ", least significant bit first" if self.reflect else ""
return (f"{self.length} bytes at bit {self.offset}{order}: {marks} "
f"{self.hex}")
def _checked(count: int, at: int, data: list[int],
reflect: bool = False) -> NearMiss:
"""Every check that framing would have to pass, run one at a time."""
if count in (7, 8, 9):
checks = (("sum", (sum(data[:-1]) & 0xFF) == data[-1]),
("parity", all(parity8(byte) == 0 for byte in data[2:-1])),
("type", (data[2] & 0x3F) in _KNOWN_TYPES))
elif count == 5:
checks = (("sum", (sum(data[:4]) & 0xFF) == data[4]),)
else:
checks = (("crc", crc8(data[:3]) == data[3]),)
family = MESSAGE_FAMILY[count]
decoder = dict(_DECODERS)[count]
return NearMiss(family=family, length=count, offset=at, data=tuple(data),
checks=checks + (("plausible", decoder(data) is not None),),
read=decoder(data) is not None, reflect=reflect)
def near_misses(bits: str, most: int = 2) -> list[NearMiss]:
"""The framings that came closest, best first.
Only the positions a real message could occupy, because a window halfway
through a burst failing its checksum says nothing anybody needs.
"""
out = []
for count, _decoder in _DECODERS:
need = count * 8
for at in range(0, len(bits) - need + 1):
if at > LEAD_BITS or len(bits) - at - need > TAIL_BITS:
continue
for reflect in BIT_ORDERS:
try:
out.append(_checked(count, at,
_bytes_at(bits, at, count, reflect),
reflect))
except ValueError:
break
out.sort(key=lambda miss: (miss.score, miss.length), reverse=True)
return out[:most]
@dataclass
class Survey:
"""What one block of samples looked like at every stage.
Only ever built to be printed. It exists because "nothing was heard" is
four different faults wearing the same coat -- no signal at all, a signal
below the gate, a burst that sliced into the wrong number of pulses, or
bits that came out and failed their checksums -- and they want four
different answers. Each stage here separates one of them from the next.
"""
quiet: float = 0.0 # the noise, as the gate measures it
gate: float = 0.0 # what a burst has to clear
peak: float = 0.0 # the loudest thing in the block
rate: float = 0.0 # what the envelope was sliced at
seen: list = field(default_factory=list) # (burst, slicings, framed)
readings: list = field(default_factory=list)
misses: list = field(default_factory=list) # per burst, the closest tries
@property
def loudest(self) -> float:
return self.peak / self.quiet if self.quiet > 0 else float("inf")
def survey(iq: np.ndarray, sample_rate: float, offset: float = 0.0,
when: float = 0.0) -> Survey:
"""One block, taken apart stage by stage, for when nothing decodes.
The readings are obtained by running the ordinary path rather than by
repeating it here, so that what this reports and what the program does
cannot come apart -- a diagnostic that disagrees with the thing it is
diagnosing is worse than none.
"""
envelope, rate = baseband(iq, sample_rate, offset)
smooth = _smoothed(envelope, rate)
peak = float(np.percentile(smooth, 99.99)) if smooth.size else 0.0
quiet = float(np.percentile(smooth, 20)) if smooth.size else 0.0
seen, misses = [], []
for burst in bursts(envelope, rate):
tries = slicings(burst)
# Without the corroboration rule, so that a message which framed
# correctly and merely arrived once is reported as what it is,
# rather than as silence.
framed = _best([r for bits in tries for r in candidates(bits)],
confirm=False)
seen.append((burst, tries, framed))
# What almost worked, for a burst that came to nothing. Judged
# across every reading of it, because which reading was the right one
# is the question and not the answer.
closest = sorted((m for bits in tries for m in near_misses(bits)),
key=lambda m: (m.score, m.length), reverse=True)
misses.append(closest[:2])
return Survey(quiet=quiet,
gate=_noise_gate(smooth, quiet, peak) if smooth.size else 0.0,
peak=peak, rate=rate, seen=seen, misses=misses,
readings=readings_from(iq, sample_rate, offset, when))
# ---------------------------------------------------------------------------
# Building the same messages, so the decoder can be held to them
# ---------------------------------------------------------------------------
#
# Every one of these lives beside the decoder that reads it, so that the two
# cannot drift apart, and so a test can put a temperature in and take the
# same one out. They are also what the simulator transmits, which means
# `--simulate` runs the real slicer and the real decoder rather than a
# shortcut past them.
def _txr_bits(data: list[int]) -> str:
"""Finish a TXR-family message: parity on the payload, then the sum."""
data = list(data)
for i in range(2, len(data)):
data[i] &= 0x7F
if parity8(data[i]) != 0: # even, as the sensors send it
data[i] |= 0x80
data.append(sum(data) & 0xFF)
return "".join(format(byte, "08b") for byte in data)
def _status(message: int, battery_low: bool) -> int:
return (message & 0x3F) | (0x00 if battery_low else 0x40)
def _wind_raw(kph: float) -> int:
return 0 if kph <= 0 else max(1, min(255, int(round((kph - 1.0) / 0.8278))))
def tower_frame(sensor: int, celsius: float, humidity: int,
channel: str = "A", battery_low: bool = False) -> str:
"""One 592TXR message: temperature, humidity, checksum and parity."""
raw = int(round((celsius + 100.0) * 10.0))
return _txr_bits([((CHANNELS.index(channel) & 3) << 6) | ((sensor >> 8) & 0x3F),
sensor & 0xFF,
_status(0x04, battery_low),
humidity & 0x7F,
(raw >> 7) & 0x0F,
raw & 0x7F])
def five_in_one_wind_rain(sensor: int, kph: float, degrees: float,
rain_counter: int, channel: str = "A",
battery_low: bool = False) -> str:
"""The 5-in-1's wind message: speed, where from, and the rain counter."""
wind = _wind_raw(kph)
point = min(range(16), key=lambda i: abs(WIND_POINTS[i] - degrees % 360))
counter = rain_counter & 0x3FFF
return _txr_bits([((CHANNELS.index(channel) & 3) << 6) | ((sensor >> 8) & 0x0F),
sensor & 0xFF,
_status(0x31, battery_low),
(wind >> 3) & 0x1F,
((wind & 0x07) << 4) | point,
(counter >> 7) & 0x7F,
counter & 0x7F])
def five_in_one_weather(sensor: int, kph: float, celsius: float,
humidity: int, channel: str = "A",
battery_low: bool = False) -> str:
"""The 5-in-1's other message: speed, temperature and humidity."""
wind = _wind_raw(kph)
raw = int(round((celsius * 9 / 5 + 32) * 10.0)) + 400
return _txr_bits([((CHANNELS.index(channel) & 3) << 6) | ((sensor >> 8) & 0x0F),
sensor & 0xFF,
_status(0x38, battery_low),
(wind >> 3) & 0x1F,
((wind & 0x07) << 4) | ((raw >> 7) & 0x0F),
raw & 0x7F,
humidity & 0x7F])
def lightning_frame(sensor: int, celsius: float, humidity: int,
strikes: int = 0, miles: int = 31,
channel: str = "A", battery_low: bool = False,
interference: bool = False) -> str:
"""One 6045M message: the weather, the strike count and how far off."""
raw = int(round(((celsius * 9 / 5 + 32) + 100.0) * 10.0))
return _txr_bits([((CHANNELS.index(channel) & 3) << 6) | ((sensor >> 8) & 0x0F),
sensor & 0xFF,
_status(LIGHTNING_MESSAGE, battery_low),
humidity & 0x7F,
(raw >> 7) & 0x1F,
raw & 0x7F,
strikes & 0x7F,
(0x20 if interference else 0x00) | (miles & 0x1F)])
def _twelve(celsius: float) -> int:
raw = int(round(celsius * 10.0))
return raw + 0x1000 if raw < 0 else raw
def frame_609(sensor: int, celsius: float, humidity: int,
battery_low: bool = False) -> str:
"""One 609TXC message: identity, temperature, humidity, a sum."""
raw = _twelve(celsius)
data = [sensor & 0xFF,
(0x00 if battery_low else 0x80) | ((raw >> 8) & 0x0F),
raw & 0xFF,
humidity & 0xFF]
data.append(sum(data) & 0xFF)
return "".join(format(byte, "08b") for byte in data)
def frame_606(sensor: int, celsius: float, battery_low: bool = False) -> str:
"""One 606TX message: identity, temperature, a CRC-8, and that is all."""
raw = _twelve(celsius)
data = [sensor & 0xFF,
(0x00 if battery_low else 0x80) | ((raw >> 8) & 0x0F),
raw & 0xFF]
data.append(crc8(data))
return "".join(format(byte, "08b") for byte in data)
# ---------------------------------------------------------------------------
# Turning bits back into something a receiver would have heard
# ---------------------------------------------------------------------------
# The timings, in microseconds. The newer sensors vary the pulse; the two
# older ones keep the pulse and vary the gap. Neither is read against these
# numbers -- the slicer compares a pulse with its own gap, or a gap with the
# other gaps -- so they are here to transmit with, not to decode by.
PWM_TIMING = {"sync_mark": 600.0, "sync_space": 600.0, "syncs": 4,
"one": (400.0, 220.0), "zero": (220.0, 400.0)}
PPM_TIMING = {"sync_mark": 500.0, "sync_space": 1000.0, "syncs": 1,
"one": (500.0, 2000.0), "zero": (500.0, 1000.0)}
TIMINGS = {"pwm": PWM_TIMING, "ppm": PPM_TIMING}
def pulse_train(bits: str, coding: str = "pwm") -> tuple[list, list]:
"""One message as the marks and spaces a transmitter would key out."""
timing = TIMINGS[coding]
marks, spaces = [], []
for _ in range(timing["syncs"]):
marks.append(timing["sync_mark"])
spaces.append(timing["sync_space"])
for bit in bits:
mark, space = timing["one" if bit == "1" else "zero"]
marks.append(mark)
spaces.append(space)
if coding == "ppm":
# In pulse-position coding the bit *is* the gap, so the last gap has
# to be closed by a pulse or there is no gap there to read. That
# trailing pulse carries no bit of its own.
marks.append(timing["sync_mark"])
return marks, spaces
def modulate(messages, sample_rate: float = 1_024_000.0,
coding: str = "pwm", repeats: int = 3, offset: float = 0.0,
amplitude: float = 1.0, noise: float = 0.0, seed: int = 0,
lead_us: float = 4_000.0, gap_us: float = 8_000.0) -> np.ndarray:
"""What a receiver would have heard, for one or more messages.
Each message is sent ``repeats`` times with a short silence between the
copies and a longer one after them, which is what the real sensors do and
is why the two thinly-checked models can be believed at all.
"""
per_us = sample_rate / 1e6
parts = [np.zeros(int(lead_us * per_us), dtype=np.float32)]
for bits in ([messages] if isinstance(messages, str) else messages):
marks, spaces = pulse_train(bits, coding)
for _ in range(max(1, repeats)):
for mark, space in zip(marks, spaces + [0.0] * len(marks)):
parts.append(np.full(int(round(mark * per_us)), amplitude,
dtype=np.float32))
if space:
parts.append(np.zeros(int(round(space * per_us)),
dtype=np.float32))
parts.append(np.zeros(int(gap_us * per_us), dtype=np.float32))
parts.append(np.zeros(int(gap_us * per_us * 2), dtype=np.float32))
envelope = np.concatenate(parts)
return _to_air(envelope, sample_rate, offset, noise, seed)
def _to_air(envelope: np.ndarray, sample_rate: float, offset: float,
noise: float, seed: int) -> np.ndarray:
"""An envelope, put where on the band the receiver expects to find it."""
signal = envelope.astype(np.complex64)
if offset:
turn = np.exp(2j * math.pi * offset
* np.arange(signal.size, dtype=np.float64) / sample_rate)
signal = signal * turn.astype(np.complex64)
if noise:
rng = np.random.default_rng(seed)
signal = signal + noise * (
rng.standard_normal(signal.size)
+ 1j * rng.standard_normal(signal.size)).astype(np.complex64)
return signal.astype(np.complex64)
# ---------------------------------------------------------------------------
# Sensors that are not there
# ---------------------------------------------------------------------------
@dataclass
class VirtualSensor:
"""A sensor on a fence post that does not exist.
It has an identity, a model, a channel switch and an opinion about the
weather, and it broadcasts all of that exactly as a real one does --
through the real encoders, the real modulator, the real slicer and the
real decoder. Nothing downstream can tell the difference, which is the
point: everything downstream is what is being exercised.
"""
family: str = "tower"
sensor: int = 0x1234
channel: str = "A"
every: float = 16.0 # seconds between transmissions
phase: float = 0.0 # so they do not all speak at once
strength: float = 1.0
celsius: float = 18.0
humidity: float = 55.0
kph: float = 8.0
rain_counter: int = 0
strikes: int = 0
storm_miles: int = 31
battery_low: bool = False
label: str = "" # what a person would call it
clock: float = 0.0
turn: int = 0
@property
def coding(self) -> str:
return "ppm" if self.family in ("609", "606") else "pwm"
def messages(self) -> list[str]:
"""What it says this time round.
The 5-in-1 alternates, because it has more to say than fits in one
message; everything else says the same thing every time.
"""
self.turn += 1
if self.family == "tower":
return [tower_frame(self.sensor, self.celsius, int(self.humidity),
self.channel, self.battery_low)]
if self.family == "5n1":
if self.turn % 2:
return [five_in_one_wind_rain(
self.sensor, self.kph, (self.turn * 23.0) % 360.0,
self.rain_counter, self.channel, self.battery_low)]
return [five_in_one_weather(self.sensor, self.kph, self.celsius,
int(self.humidity), self.channel,
self.battery_low)]
if self.family == "6045":
return [lightning_frame(self.sensor, self.celsius,
int(self.humidity), self.strikes,
self.storm_miles, self.channel,
self.battery_low)]
if self.family == "609":
return [frame_609(self.sensor, self.celsius, int(self.humidity),
self.battery_low)]
return [frame_606(self.sensor, self.celsius, self.battery_low)]
def advance(self, seconds: float, rng) -> None:
"""Let the weather move on, slowly and within reason.
Slowly matters. A garden that swings six degrees in a minute would
exercise everything downstream perfectly well and would look, to
anyone running ``--simulate`` to see what the program does, like a
program that cannot read a thermometer. So the rates here are
roughly a real afternoon's: about a degree an hour of drift, wind
that gusts around a mean, and rain that falls in tenths of a
millimetre at a time. Everything stays inside what the hardware
could report, so the decoder's own plausibility check is never the
thing that rejects it.
"""
self.clock += seconds
turn = math.sin(self.clock / 1800.0) # half an hour a swing
self.celsius = min(45.0, max(-30.0, self.celsius
+ seconds * (0.0004 * turn
+ rng.normal(0.0, 0.0008))))
self.humidity = min(99.0, max(5.0, self.humidity
+ seconds * (-0.0016 * turn
+ rng.normal(0.0, 0.004))))
gust = 1.0 + 0.5 * math.sin(self.clock / 41.0)
self.kph = min(90.0, max(0.0, self.kph
+ seconds * 0.08 * (gust * 11.0 - self.kph)))
if self.family == "5n1" and rng.random() < 0.01 * seconds:
self.rain_counter += 1
if self.family == "6045":
if rng.random() < 0.006 * seconds:
self.strikes = (self.strikes + 1) & 0x7F
self.storm_miles = int(rng.integers(1, 25))
elif rng.random() < 0.002 * seconds:
self.storm_miles = 31
def default_sensors() -> list[VirtualSensor]:
"""One of each, in a garden that does not exist.
The transmission periods are deliberately awkward numbers. Two sensors
on exactly sixteen seconds would transmit over the top of each other
every time they lined up and never come unstuck again, which would be
both unrealistic and unfair: real ones are timed by a cheap oscillator in
the cold and drift past each other within a few minutes. Here they are
simply given periods with no common factor, which has the same effect and
keeps the garden reproducible from a seed.
"""
return [
VirtualSensor(family="tower", sensor=0x1A2B, channel="A", every=16.1,
phase=1.0, celsius=17.4, humidity=62,
label="back fence"),
VirtualSensor(family="tower", sensor=0x0C41, channel="B", every=16.7,
phase=6.5, celsius=20.9, humidity=44, strength=0.55,
label="greenhouse"),
VirtualSensor(family="5n1", sensor=0x0777, channel="A", every=18.3,
phase=3.0, celsius=16.8, humidity=71, kph=11.0,
rain_counter=1284, label="mast"),
VirtualSensor(family="6045", sensor=0x0311, channel="C", every=24.9,
phase=9.0, celsius=19.1, humidity=58, strikes=3,
storm_miles=12, label="lightning"),
VirtualSensor(family="609", sensor=0x5C, every=30.2, phase=12.0,
celsius=4.2, humidity=80, battery_low=True,
label="shed"),
VirtualSensor(family="606", sensor=0x93, every=32.6, phase=15.0,
celsius=-3.5, strength=0.7, label="freezer"),
]
class SimulatedSensors:
"""A receiver-shaped source of weather that is not happening.
It answers ``read_samples`` the way the dongle does and hands back the
same on-off-keyed bursts a garden full of sensors would, so the whole
section -- the slicer, every decoder, the naming, the log, the report and
the export -- can be run through without an aerial, a sensor or a garden.
"""
def __init__(self, sensors=None, sample_rate: float = 1_024_000.0,
offset: float = 250_000.0, noise: float = 0.02,
seed: int = 0, realtime: bool = False):
self.sensors = list(sensors) if sensors is not None \
else default_sensors()
self.sample_rate = float(sample_rate)
self.offset = float(offset)
self.noise = noise
self.rng = np.random.default_rng(seed)
self.frequency = ACURITE_HZ
self.clock = 0.0
# A block is a second of samples but takes longer than a second to
# slice, so sensors advanced by the length of the block fall behind
# the clock their readings are stamped with. Listening for real, the
# weather moves by the time that actually passed; in a test, by the
# block, so the same seed gives the same garden twice.
self.realtime = realtime
self._last = None
# -- the shape of a device -------------------------------------------
def open(self):
return self
def close(self) -> None:
return None
def tune(self, hz: float, settle: bool = True) -> int:
self.frequency = hz
return int(hz)
def read_samples(self, count: int, flush: bool = False) -> np.ndarray:
"""One block of garden: whoever is due to speak, speaks."""
import time as _time
seconds = count / self.sample_rate
if self.realtime:
# A dongle hands back a second of samples once a second has
# passed, and this stands in for a dongle, so it waits. Without
# the wait a block comes back the moment it is asked for, the
# garden lives several times faster than the clock the readings
# are stamped with, and anything measured against wall time --
# how long to listen for, how large a capture will be -- comes
# out wrong by whatever the machine happens to be worth.
now = _time.monotonic()
if self._last is not None:
behind = seconds - (now - self._last)
if behind > 0:
_time.sleep(behind)
now = _time.monotonic()
seconds = max(0.0, now - self._last)
self._last = now
block = np.zeros(count, dtype=np.complex64)
began, ended = self.clock, self.clock + seconds
for sensor in self.sensors:
for at in _due(sensor, began, ended):
burst = modulate(sensor.messages(), self.sample_rate,
coding=sensor.coding, offset=self.offset,
amplitude=sensor.strength, lead_us=0.0)
start = int((at - began) * self.sample_rate)
room = min(burst.size, max(0, count - start))
if room > 0:
block[start:start + room] += burst[:room]
sensor.advance(seconds, self.rng)
self.clock = ended
if self.noise:
block = block + self.noise * (
self.rng.standard_normal(count)
+ 1j * self.rng.standard_normal(count)).astype(np.complex64)
return block.astype(np.complex64)
def read_seconds(self, seconds: float, flush: bool = False) -> np.ndarray:
return self.read_samples(int(self.sample_rate * seconds))
def _due(sensor: VirtualSensor, began: float, ended: float) -> list[float]:
"""Every moment in a block at which one sensor is due to transmit."""
if sensor.every <= 0 or ended <= began:
return []
first = math.ceil((began - sensor.phase) / sensor.every)
out = []
at = sensor.phase + first * sensor.every
while at < ended:
if at >= began:
out.append(at)
at += sensor.every
return out