Most of what identifies itself on the air identifies itself in Morse. A repeater, a beacon, an unattended transmitter: four to six characters, over in a second or two, and no speech anywhere in the capture. Every one of those was being thrown away, in three separate places. The callsign book and the map were built inside the transcription branch, on the reasoning that callsigns come out of transcripts. They also come out of Morse and out of APRS headers, neither of which involves a speech recogniser -- so a receiver with none installed found none of them, and a CW ident reached the sidecar and stopped there. Both are now built whenever classification is on, and all three sources go through one place. The CW decoder only ran where the classifier had already said cw, ook or carrier. A two-second ident is a fraction of a capture named after whatever filled the rest of it. Every capture is offered to it now, once it has finished; a decode does not relabel a capture that plainly holds speech. And the decoder's own gates were written for a paragraph. Three characters, eight elements, and any repeated character refused -- which read VVV, DE, AR and K correctly and then discarded them. Short is the normal case now, on a second bar: perfect timing, nothing undecoded, and the keyed tone at least 20 dB over its band. That last is not decoration. With four elements the dot length is fitted to those very elements, so noise lands on the grid as neatly as keying does; a third of a second of white noise decodes as a perfectly timed V. Over 200 noise blocks the loudest bin never rose 13 dB above the median while keying at 3 dB SNR sits above 40. One keyed element is still refused: a single pulse is an E or a T whether a person sent it or the squelch opened on a click. 288 non-Morse cases, no false positives. Feeding that text to a callsign lookup made truncation matter. A capture opens when the squelch does, halfway through an element as often as not, and half a character is not a smaller reading -- a K missing its first dash is an A. So the sliced character is dropped, and so is the rest of its word, because what is left can read as a whole one: K1AA caught halfway through is K1A, which is somebody else. Across 1805 truncated captures that is 107 invented callsigns down to none, with 550 correct ones still found. Phonetics, which is the other half of the ask. A recogniser has never heard of the alphabet -- it writes what the words sounded like: Whiskey-One-Alpha-Whiskey hyphenated WhiskeyOneAlphaWhiskey run together Whiskey1AlphaWhiskey and half in digits wiskey one alfa whisky spelled the way it sounded whiskey one alpha, uh, whiskey with the hesitation written down All read back to W1AW now. A word is only taken apart when it is phonetic all the way through, which is what keeps it off "kilometre" and "victorious". Two bugs found on the way, both of which invented a callsign: - Nothing is joined across a slash any more. The beacon W1AW/B came back as W1AWB, which belongs to nobody, and W1AW-4 came back as nothing at all. - Nor across a gap the sender chose. A transcript's spacing is the recogniser's guess and may be closed up; a word gap in Morse is seven dot units, so "KU0W K" is a station signing off, not a longer callsign. Also here: - saunterbrowse gives Morse a panel of its own, with the licence under it, searchable with / and readable with t. - --simulate no longer looks anything up or writes a map. The demo band is invented but W1AW is the ARRL's own station, and it would have been pinned to the same map a real scan writes. - classify._psk_order took the logarithm of zero on a silent block. - The demo band has a repeater ident in it, and its Morse no longer runs one repeat into the next. - conftest refuses a real licence lookup from any test. 1161 tests, up from 1014. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PsWPTweCT6pwxKngvVxcg
400 lines
17 KiB
Python
400 lines
17 KiB
Python
"""Deciding whether a capture actually carries a signal worth keeping.
|
|
|
|
Power alone cannot tell a transmission from a hump of interference, so every
|
|
capture is assessed for *content* before it is kept:
|
|
|
|
* **voice** -- speech structure in the demodulated audio
|
|
* **cw** -- a keyed carrier whose timing resolves as Morse
|
|
* **digital** -- a symbol rate, discrete FSK levels, or an M-PSK phase line
|
|
* **trunk** -- a trunking control channel: continuous data, never any speech
|
|
* **carrier** -- a steady unmodulated carrier (real, but carries nothing)
|
|
* **noise** -- no structure at all: static, interference, receiver artefacts
|
|
|
|
Only the categories the user asked for are recorded.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import math
|
|
from dataclasses import dataclass, field
|
|
|
|
import numpy as np
|
|
|
|
__all__ = ["VoiceMetrics", "voice_metrics", "Assessment", "assess",
|
|
"CATEGORIES", "RAYLEIGH_CV"]
|
|
|
|
CATEGORIES = ("voice", "cw", "digital", "trunk", "carrier", "noise")
|
|
|
|
# Envelope coefficient of variation for complex Gaussian noise: |x| is
|
|
# Rayleigh distributed, so std/mean is exactly sqrt(4/pi - 1).
|
|
RAYLEIGH_CV = math.sqrt(4.0 / math.pi - 1.0) # 0.5227
|
|
|
|
|
|
@dataclass
|
|
class VoiceMetrics:
|
|
"""Speech-likeness measurements taken on demodulated audio."""
|
|
|
|
voiced_fraction: float = 0.0 # frames with a clear pitch period
|
|
voiced_frames: int = 0 # how much evidence that fraction rests on
|
|
pitch_hz: float = 0.0 # median pitch of the voiced frames
|
|
pitch_variation: float = 0.0 # speech pitch drifts; a buzz does not
|
|
dynamic_range_db: float = 0.0 # speech pauses between phrases
|
|
syllabic: float = 0.0 # envelope modulation in the 2-8 Hz band
|
|
band_concentration: float = 0.0 # energy inside 250-3400 Hz
|
|
spectral_flux: float = 0.0 # formants moving between frames
|
|
active_fraction: float = 0.0 # how much of the clip is above silence
|
|
score: float = 0.0
|
|
|
|
def describe(self) -> str:
|
|
return (f"voiced {self.voiced_fraction:.2f}, pitch "
|
|
f"{self.pitch_hz:.0f} Hz +-{self.pitch_variation*100:.0f}%, "
|
|
f"dynamics {self.dynamic_range_db:.0f} dB, "
|
|
f"syllabic {self.syllabic:.2f}")
|
|
|
|
|
|
def _frames(x: np.ndarray, fs: float, win_s: float = 0.032,
|
|
hop_s: float = 0.010):
|
|
n = int(win_s * fs)
|
|
hop = max(1, int(hop_s * fs))
|
|
if x.size < n * 3:
|
|
return np.zeros((0, n)), hop
|
|
count = 1 + (x.size - n) // hop
|
|
idx = np.arange(n)[None, :] + hop * np.arange(count)[:, None]
|
|
return x[idx], hop
|
|
|
|
|
|
def _pitch_track(frames: np.ndarray, fs: float,
|
|
f_lo: float = 70.0, f_hi: float = 400.0):
|
|
"""Normalised autocorrelation peak per frame, over the human pitch range.
|
|
|
|
Computed for every frame at once through the FFT. A direct
|
|
autocorrelation is quadratic in the frame length, and running it frame by
|
|
frame was slow enough to stall the capture loop -- which costs samples,
|
|
and shows up as a recording that plays too fast.
|
|
|
|
Returns ``(strengths, pitches_hz)``.
|
|
"""
|
|
n = frames.shape[1]
|
|
if frames.shape[0] == 0 or n < 16:
|
|
return np.zeros(0), np.zeros(0)
|
|
f = frames - frames.mean(axis=1, keepdims=True)
|
|
energy = np.einsum("ij,ij->i", f, f)
|
|
|
|
nfft = 1 << int(math.ceil(math.log2(2 * n)))
|
|
spec = np.fft.rfft(f, nfft, axis=1)
|
|
ac = np.fft.irfft(spec * np.conj(spec), nfft, axis=1)[:, :n]
|
|
|
|
lo = max(1, int(fs / f_hi))
|
|
hi = min(n - 1, int(fs / f_lo))
|
|
if hi <= lo + 4:
|
|
return np.zeros(frames.shape[0]), np.zeros(frames.shape[0])
|
|
|
|
seg = ac[:, lo:hi]
|
|
k = np.argmax(seg, axis=1)
|
|
peak = seg[np.arange(seg.shape[0]), k]
|
|
strength = np.where(energy > 1e-12, peak / np.maximum(energy, 1e-12), 0.0)
|
|
# A maximum sitting on the edge of the search range is the search running
|
|
# out of room, not a periodicity. Noise does this constantly, and the
|
|
# result is a stream of "pitches" pinned at exactly the limit.
|
|
edge = (k <= 1) | (k >= seg.shape[1] - 2)
|
|
strength = np.where(edge, 0.0, strength)
|
|
pitch = np.where(edge, 0.0, fs / (lo + k))
|
|
return strength, pitch
|
|
|
|
|
|
def voice_metrics(audio: np.ndarray, fs: float) -> VoiceMetrics:
|
|
"""Measure how speech-like a block of audio is.
|
|
|
|
Speech has four properties that static, hum, tones and data bursts do not
|
|
have together: a pitch period in the 70-400 Hz range during voiced sounds,
|
|
pauses between phrases, an envelope that varies at the syllable rate, and
|
|
formants that move from frame to frame.
|
|
"""
|
|
m = VoiceMetrics()
|
|
audio = np.asarray(audio, dtype=np.float64).ravel()
|
|
if audio.size < int(fs * 0.4):
|
|
return m
|
|
audio = audio - audio.mean()
|
|
peak = float(np.abs(audio).max())
|
|
if peak < 1e-6:
|
|
return m
|
|
audio = audio / peak
|
|
|
|
frames, hop = _frames(audio, fs)
|
|
if frames.shape[0] < 16:
|
|
return m
|
|
win = np.hanning(frames.shape[1])
|
|
|
|
energy = np.sqrt(np.mean((frames * win) ** 2, axis=1)) + 1e-12
|
|
e_db = 20.0 * np.log10(energy)
|
|
|
|
# Speech pauses show up as a wide spread between loud and quiet frames.
|
|
m.dynamic_range_db = float(np.percentile(e_db, 95) - np.percentile(e_db, 15))
|
|
|
|
# Only look for pitch where there is something to look at.
|
|
gate = e_db > (np.percentile(e_db, 95) - 25.0)
|
|
m.active_fraction = float(gate.mean())
|
|
|
|
idx = np.flatnonzero(gate)[:400]
|
|
if idx.size:
|
|
strengths, pitches = _pitch_track(frames[idx], fs)
|
|
voiced = strengths > 0.38
|
|
m.voiced_fraction = float(voiced.mean())
|
|
m.voiced_frames = int(voiced.sum())
|
|
if np.any(voiced):
|
|
track = np.asarray(pitches)[voiced]
|
|
m.pitch_hz = float(np.median(track))
|
|
if m.pitch_hz > 0 and track.size > 3:
|
|
# Intonation: a talker's pitch wanders, mains hum and a
|
|
# switching-supply buzz sit on exactly one frequency.
|
|
m.pitch_variation = float(np.std(track) / m.pitch_hz)
|
|
|
|
# Syllable-rate envelope modulation, 2-8 Hz.
|
|
frame_rate = fs / hop
|
|
env = energy - energy.mean()
|
|
if env.size >= 64 and frame_rate > 24:
|
|
n = 1 << int(math.floor(math.log2(env.size)))
|
|
spec = np.abs(np.fft.rfft(env[:n] * np.hanning(n), n))
|
|
mf = np.fft.rfftfreq(n, 1.0 / frame_rate)
|
|
syl = (mf >= 2.0) & (mf <= 8.0)
|
|
ref = (mf >= 0.5) & (mf <= 20.0)
|
|
if np.any(syl) and np.any(ref):
|
|
total = float(np.sum(spec[ref] ** 2))
|
|
if total > 0:
|
|
m.syllabic = float(np.sum(spec[syl] ** 2) / total)
|
|
|
|
# Energy inside the voice band, and how much the spectrum moves.
|
|
mag = np.abs(np.fft.rfft(frames * win, axis=1))
|
|
freqs = np.fft.rfftfreq(frames.shape[1], 1.0 / fs)
|
|
band = (freqs >= 250.0) & (freqs <= 3400.0)
|
|
tot = np.sum(mag ** 2) + 1e-12
|
|
m.band_concentration = float(np.sum(mag[:, band] ** 2) / tot)
|
|
|
|
norm = mag / (np.linalg.norm(mag, axis=1, keepdims=True) + 1e-12)
|
|
if norm.shape[0] > 1:
|
|
m.spectral_flux = float(np.mean(np.linalg.norm(np.diff(norm, axis=0),
|
|
axis=1)))
|
|
|
|
# ---- combine -------------------------------------------------------
|
|
# A pitch track that moves is the one measurement here that only a voice
|
|
# produces, so it gates the score rather than contributing a share of it.
|
|
# Dynamics, syllable-rate modulation and energy landing in the voice band
|
|
# are all things that static does too: weighted alongside voicing they
|
|
# were enough to carry noise over the line on their own, which is exactly
|
|
# how hiss ended up being recorded as speech.
|
|
core = min(1.0, m.voiced_fraction / 0.30)
|
|
# Pitch drift is the guard against tones and hum, but measuring drift
|
|
# needs several voiced frames to measure it across. A short over -- a
|
|
# two-second "QSL, 73" -- cannot supply them, so the requirement eases
|
|
# when the evidence is thin; the steady-tone guard below still applies.
|
|
if m.voiced_frames >= 20:
|
|
core *= min(1.0, m.pitch_variation / 0.04)
|
|
else:
|
|
core *= min(1.0, 0.45 + m.pitch_variation / 0.04)
|
|
|
|
dyn_term = min(1.0, max(0.0, (m.dynamic_range_db - 6.0) / 20.0))
|
|
syl_term = min(1.0, max(0.0, (m.syllabic - 0.18) / 0.30))
|
|
band_term = min(1.0, max(0.0, (m.band_concentration - 0.25) / 0.45))
|
|
support = 0.40 * dyn_term + 0.35 * syl_term + 0.25 * band_term
|
|
|
|
score = core * (0.55 + 0.45 * support)
|
|
|
|
# A steady tone is perfectly periodic and would score full marks on pitch;
|
|
# what it does not have is a spectrum that changes or any dynamics.
|
|
if m.spectral_flux < 0.045 and m.dynamic_range_db < 6.0:
|
|
score *= 0.25
|
|
m.score = float(min(1.0, score))
|
|
return m
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
|
|
@dataclass
|
|
class Assessment:
|
|
category: str = "noise"
|
|
score: float = 0.0
|
|
accept: bool = False
|
|
reason: str = ""
|
|
voice: VoiceMetrics = field(default_factory=VoiceMetrics)
|
|
noise_likeness: float = 0.0
|
|
control: object = None # ControlChannel when this is a control channel
|
|
|
|
def describe(self) -> str:
|
|
return f"{self.category} ({self.score:.2f}): {self.reason}"
|
|
|
|
|
|
def noise_likeness(f) -> float:
|
|
"""0 = structured signal, 1 = indistinguishable from receiver noise.
|
|
|
|
Gaussian noise has a known envelope statistic, a flat spectrum, no
|
|
carrier, no discrete frequency levels and no symbol rate. Scoring all of
|
|
those together is far more reliable than any one of them.
|
|
"""
|
|
votes = []
|
|
# Envelope statistics sitting on the Rayleigh value.
|
|
votes.append(max(0.0, 1.0 - abs(f.env_cv - RAYLEIGH_CV) / 0.16))
|
|
# Flat, featureless spectrum.
|
|
votes.append(min(1.0, max(0.0, (f.flatness - 0.35) / 0.45)))
|
|
# No carrier line worth the name.
|
|
votes.append(min(1.0, max(0.0, (14.0 - f.papr_spectral) / 10.0)))
|
|
# No symbol rate.
|
|
votes.append(1.0 if f.baud_strength < 11.0 else 0.0)
|
|
# No discrete discriminator levels.
|
|
votes.append(1.0 if f.freq_modes <= 1 else 0.0)
|
|
# No keying.
|
|
votes.append(min(1.0, max(0.0, (9.0 - f.ook_contrast_db) / 9.0)))
|
|
return float(np.mean(votes))
|
|
|
|
|
|
def digital_structure(f) -> tuple[float, list[str]]:
|
|
"""Score how much symbol structure a signal shows, with the evidence.
|
|
|
|
A symbol *rate* on its own proves nothing -- the cyclostationary estimator
|
|
will always return its best peak, and noise and bare carriers both produce
|
|
one. So an identified keying scheme is required first, and the baud
|
|
figure only corroborates it.
|
|
"""
|
|
score = 0.0
|
|
bits: list[str] = []
|
|
|
|
if f.freq_modes >= 2 and f.mode_spacing > 0 and f.level_dwell > 0.32:
|
|
score += 0.45
|
|
bits.append(f"{f.freq_modes} discrete frequency levels "
|
|
f"{f.mode_spacing:.0f} Hz apart")
|
|
# A phase line needs a symbol rate behind it that holds still; without
|
|
# one, an Mth-power peak is just the strongest thing in a noisy spectrum.
|
|
stable_baud = (f.baud_strength > 14 and 50.0 <= f.baud <= 100_000.0
|
|
and f.baud_stability > 0.80)
|
|
if f.psk_strength > 22 and f.psk_order and f.ifreq_kurtosis > 3.0 \
|
|
and f.psk_strength > f.papr_spectral + 8.0 \
|
|
and f.carrier_ratio < 0.10 and stable_baud:
|
|
score += 0.40
|
|
bits.append(f"{f.psk_order}-PSK phase line")
|
|
# On/off contrast on its own is not evidence of keying: a signal fading
|
|
# in and out at the squelch edge produces plenty of it. Real OOK data
|
|
# also has a symbol rate, so require both.
|
|
plausible_baud = stable_baud
|
|
# Keying means the on and off runs land on a common symbol grid. Contrast
|
|
# and a baud estimate are not enough on their own: a signal fading across
|
|
# the squelch produces both, with run lengths that fit no grid at all.
|
|
if (f.ook_contrast_db > 16 and 0.05 < f.ook_duty < 0.95
|
|
and plausible_baud and f.keying_regularity > 0.55):
|
|
score += 0.35
|
|
bits.append(f"on/off keying, {f.ook_contrast_db:.0f} dB contrast, "
|
|
f"runs on a {f.keying_regularity*100:.0f}% regular grid")
|
|
|
|
if score <= 0:
|
|
return 0.0, []
|
|
|
|
if plausible_baud:
|
|
score += 0.20
|
|
bits.append(f"{f.baud:.0f} baud symbol rate")
|
|
return score, bits
|
|
|
|
|
|
def assess(classification, audio: np.ndarray, audio_rate: float,
|
|
morse=None, min_voice: float = 0.45, freq_hz: float = 0.0,
|
|
accept: tuple[str, ...] = ("voice", "cw", "digital"),
|
|
min_score: float = 0.45, skip_control: bool = True) -> Assessment:
|
|
"""Decide what a capture contains and whether it should be kept.
|
|
|
|
Deliberately does not route on the classifier's label. The label is a
|
|
best guess that can be wrong -- speech on a quiet FM channel can look like
|
|
two-level FSK, for instance -- and a capture should be kept or dropped on
|
|
what is measurably *in* it, not on what it was called.
|
|
"""
|
|
f = classification.features
|
|
a = Assessment()
|
|
if f is None:
|
|
a.reason = "no measurements available"
|
|
return a
|
|
|
|
a.noise_likeness = noise_likeness(f)
|
|
fam = classification.family
|
|
vm = voice_metrics(audio, audio_rate)
|
|
a.voice = vm
|
|
dig_score, dig_bits = digital_structure(f)
|
|
|
|
# 1. A trunking control channel, before anything else. It is a positive
|
|
# identification rather than a failure to find content -- the stream is
|
|
# real data, and would otherwise be kept as "digital" and hold the
|
|
# receiver for the full record time on a channel with nothing to hear.
|
|
if skip_control and getattr(classification, "control", None) is not None:
|
|
c = classification.control
|
|
a.category = "trunk"
|
|
a.control = c
|
|
a.score = float(c.confidence)
|
|
a.reason = f"{c.describe()}: {c.reason}"
|
|
|
|
# 2. Morse that actually decoded.
|
|
#
|
|
# Every capture is offered to the CW decoder now, not only the ones
|
|
# that looked keyed, because a station identifying itself in Morse
|
|
# sends a burst of a second or two and the classifier names the
|
|
# capture after whatever fills the rest of it. That is worth
|
|
# decoding, and it is not worth relabelling a conversation over: where
|
|
# the modulation was never keyed and there is speech in the audio, the
|
|
# speech is what the capture is, and the Morse text is recorded beside
|
|
# it either way.
|
|
elif (morse is not None and morse.is_morse
|
|
and (fam in ("cw", "ook", "carrier") or vm.score < min_voice)):
|
|
a.category = "cw"
|
|
a.score = float(morse.confidence)
|
|
a.reason = f'Morse decoded at {morse.wpm:.0f} WPM: "{morse.text.strip()[:40]}"'
|
|
|
|
# 3. Speech, whatever the modulation was called.
|
|
elif vm.score >= min_voice:
|
|
a.category = "voice"
|
|
a.score = vm.score
|
|
a.reason = f"speech in the audio ({vm.describe()})"
|
|
|
|
# 4. Symbol structure.
|
|
elif dig_score > 0.3:
|
|
a.category = "digital"
|
|
a.score = min(0.95, 0.35 + dig_score)
|
|
a.reason = "digital modulation: " + ", ".join(dig_bits)
|
|
|
|
# 5. A keyed carrier whose timing would not resolve as Morse.
|
|
elif fam == "cw" and f.ook_contrast_db > 12:
|
|
a.category = "cw"
|
|
a.score = 0.55
|
|
a.reason = (f"keyed carrier, {f.ook_contrast_db:.0f} dB on/off contrast "
|
|
"(timing did not resolve as Morse)")
|
|
|
|
# 6. Broadcast FM gets a lenient path because most of it is music, which
|
|
# has no speech pitch track to find. That leniency is confined to
|
|
# signals that really are broadcast: in the FM band, or carrying a
|
|
# 19 kHz stereo pilot. Without the restriction, any wideband hump --
|
|
# a clock harmonic, say -- walks straight through it.
|
|
elif (f.stereo_pilot or 87.5e6 <= freq_hz <= 108.1e6) \
|
|
and f.bandwidth > 50_000 and f.fdev_rms > 2_000 \
|
|
and (vm.spectral_flux > 0.15 or vm.dynamic_range_db > 8):
|
|
a.category = "voice"
|
|
a.score = max(0.55, vm.score)
|
|
a.reason = ("programme audio on a wideband FM carrier"
|
|
+ (" with a 19 kHz stereo pilot" if f.stereo_pilot else ""))
|
|
|
|
# 7. A carrier that is really just a carrier.
|
|
elif fam == "carrier" or (f.am_depth < 0.004 and f.fdev_rms < 150
|
|
and f.ook_contrast_db < 6):
|
|
a.category = "carrier"
|
|
a.score = 0.6
|
|
a.reason = "steady unmodulated carrier"
|
|
|
|
else:
|
|
a.category = "noise"
|
|
a.score = vm.score
|
|
a.reason = (f"no speech, no symbol structure and no keying "
|
|
f"({vm.describe()})")
|
|
|
|
# A strong noise verdict overrides any label, however confident. A
|
|
# score bar here left a hole: a single piece of "digital" evidence scored
|
|
# 0.80 and sailed past a check that only applied below 0.75.
|
|
if a.category != "noise" and a.noise_likeness > 0.82:
|
|
a.category = "noise"
|
|
a.reason = (f"statistics match receiver noise "
|
|
f"(noise likeness {a.noise_likeness:.2f})")
|
|
a.score = 0.0
|
|
|
|
a.accept = a.category in accept and a.score >= min_score
|
|
return a
|