MIUI / HyperOS hands out a zip of app logs, ANR traces and tcpdump captures
with the real bugreport-<device>-<timestamp>.zip nested inside. The outer
archive carries none of the entry points the bug report modules read, so every
module reported it found no files and check-bugreport still exited 0 with an
empty result — an empty analysis that looks like a finished one.
Name the entry points once in modules/bugreport/base.py, next to the
_get_dumpstate_file() that tries them, and when an archive has none of them,
open its zip members in memory and use the first one that does.
Measured over 44 bug report collections: the 14 wrapper collections go from 0
artifacts to 260 artifacts and 586638 records; the 30 normal collections are
unchanged, the descent being unreachable for them.
Fixes#935
Two issues from review:
* The set of printed state sections was global, so a section printed for
one user decided the flags of another. A user with an installed list and
no enabled section was marked enabled=False and got a LOW "installed, not
enabled" alert. Sections are now tracked per user.
* A count-only record was added only when a user had no named component at
all. A dump stating installedServiceCount=2 and naming one service read
as complete. The count-only record is now added whenever the stated count
exceeds the distinct named components for that user, and carries the
difference in a new unnamed_service_count field. The module summary adds
up that remainder.
Most builds never print the `installed services: {...}` block. They state
installedServiceCount=N in the user's attributes line and list nothing, so the
parser recorded no service at all and MVT logged "Identified a total of 0
accessibility services" — an artifact that says "no accessibility services"
about a dump that said there are five. The dump pasted in #744 is itself an
example: it states installedServiceCount=6 and names none.
Parse the count per user and, for a user whose services the dump did not list,
keep one record carrying it, raised as LOW: a count is a coverage statement,
not a running service.
Name the service state in the alert and split its severity, as asked on #744:
a service the dump states is switched off is LOW, enabled or bound is MEDIUM,
and a dump that does not state the enabled state stays MEDIUM, because "not
stated" is not "not enabled". A state section the dump never printed now reads
None rather than False. IOC matching is unaffected by the state: a disabled
service still returns CRITICAL on a match.
Measured over 163 bug reports carrying an accessibility section: MEDIUM alerts
305 -> 4, LOW 0 -> 407. The four that stay MEDIUM are the only records in the
set with enabled=True.
Fixes#744
extract_command_section() ends a section at any line starting with "------".
dumpstate prints a section's timing line when THAT section finishes, and it can
land in the middle of the section currently being written, so the section is
truncated at an arbitrary point and the rest is silently dropped.
Skip the timing line instead of returning it: it never reaches a parser as
content, and the real "------ <NAME> ------" boundary still ends the section.
Measured over 40 bug reports carrying a dumpstate: SYSTEM PROPERTIES was
truncated on 6 of them, losing 6559 properties, and 3 parsed no property at
all. On one Samsung archive getprop goes from 418 to 1259 properties and from
0 to 473 ro.* ones, bringing back ro.product.model,
ro.build.version.security_patch and ro.boot.verifiedbootstate. The other 34
archives are byte-identical.
Fixes#938
/proc/PID/mountinfo carries two option sets with different meaning: fields[5]
are the per-mount (VFS) flags, the field after the "-" separator belongs to the
superblock. A mount is writable only if both allow it.
parse_mountinfo() merged both into one list and set
is_read_write = "rw" in options, so a "rw" VFS mount over a read-only
superblock was reported as writable. On stock Xiaomi-family builds that made
every read-only mi_ext customisation overlay a HIGH "system partition is
mounted as read-write".
Require "rw" in both layers, and let "rw" count as a suspicious mount option
only when the mount is actually writable; remount, noatime and nodiratime keep
their current meaning in either layer. mount_options and options_list still
carry both layers, so nothing downstream loses data.
Measured on 50 bug reports carrying mountinfo - the 28 where the rule changes
the output plus 22 controls, 14 brands, Android 10-16: HIGH 94 -> 0,
MEDIUM 119 -> 25, the 22 controls identical, and 36233 mount entries parsed
either way. Every removed alert is a read-only superblock under a "rw" VFS
mount; no report gains an alert. A partition that really is writable still
raises the HIGH, which the new test asserts explicitly.
Fixes#936
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The tests read the settings file of whoever runs them, and importing
mvt.common.config writes it back. On a machine whose config.yaml sets
NETWORK_ACCESS_ALLOWED to false, the plugin update and URL batch tests
fail before their mocked requests are reached, and every Command run in
the suite parses the indicators downloaded on that machine, which made
the suite take minutes instead of seconds.
MVT_CONFIG_FOLDER and MVT_DATA_FOLDER in the environment now relocate
the settings file and the downloaded indicators with their update-check
state. The test conftest points both at a throwaway folder before any
mvt module is imported, and removes it at exit. Subprocesses started by
the tests inherit the variables; the isolated interpreter helper already
gives them a temporary home.
On the machine that prompted this the suite goes from 6 failures in six
minutes to none in seven seconds, and config.yaml is left alone.
Device-generated sysdiagnose archives carry a ._name entry beside every
file that has extended attributes, an ACL or Finder info; one iOS 26
archive held 1234 of them among 3648 members, and the count grows with
each release. bsdtar folds them back into the file on extraction and
hides them from listings, but tarfile returns them as regular members,
so check-sysdiagnose extracted them and handed them to every module.
A module that globs for plists or logs then tries to parse AppleDouble
headers and logs one warning per sidecar.
Leave them out of the file list, both for archives and for folders
extracted on a system that keeps them as files.
* Defer indicator loading until first use
* Load indicators once at the start of a run
The lazy property loads indicators wherever they are first read, which
in Command.run() is inside the module loop, after init(). For
check-androidqf that means the whole acquisition is walked before a
missing --iocs file is reported, and the STIX parsing lines land between
the module list and the first module.
Read the indicators once after the module list is settled, so that a bad
--module name still loads nothing, and hand the same object to every
module. check-iocs reads them once too, so that a wrong --iocs path is
reported even when no stored result matches a module.
* Drop two test assertions the change does not need
Retrying after a failed indicator load is incidental to the property
rather than a requirement, so nothing should pin it. The exact wording
of the check-backup rejection belongs to that command's own tests; exit
code 1 already shows the path was rejected before indicators were
loaded.
* Create the output folder and command.log when a run starts
Command.__init__ created the --output folder and attached the
command.log handler, so `--list-modules -o out` or a rejected target
path left an empty folder behind for a run that never happened.
Do both at the top of run(), before the module list is resolved so that
its warnings still reach the log. The nested commands run by
check-androidqf re-attach the handler at the start of their runs, as
they did at construction. Lines the CLI logs between construction and
run(), such as the target path being checked, no longer reach
command.log; info.json records the target path.
This is the remaining part of #888.
* Cache the indicators with functools.cached_property
A property with a setter and a backing attribute does by hand what
cached_property does: compute on first read, store the result on the
instance, and accept assignment, which is how nested commands receive
their parent's indicators. A failed load is still not cached.
* Announce the target from the command so that it reaches command.log
The CLI logged which backup, filesystem or acquisition it was about to
check just before run(), which is now before command.log exists, so the
line only reached the console. Each command logs it from init() instead,
which runs once the log is attached, and run() logs the module list
after init() so that the target still comes first. Nested commands
without a target path log nothing, as before.
* Log the Android backup path from the typed local
mypy cannot determine the type of target_path in this command, which the
surrounding lines already silence, so read the announced path from the
local that carries the type instead of adding another ignore.
---------
Co-authored-by: bitmeta69 <206962233+bitmeta69@users.noreply.github.com>
Co-authored-by: besendorf <janik@besendorf.org>
Co-authored-by: Donncha Ó Cearbhaill <donncha.ocearbhaill@amnesty.org>
A bugreport zip stores each entry's time as the device's wall clock, and
an unpacked bugreport's mtimes are whatever the extraction left. The
bugreport modules read both as naive local datetimes, so a tombstone's
file_timestamp moved with the analysing machine's timezone and sat three
hours from the crash time the tombstone itself records for a device in
Nairobi.
check-bugreport now resolves the device's timezone once, from --timezone
or from persist.sys.timezone in the dumpstate's SYSTEM PROPERTIES, and
hands it to the modules as device_timezone the way check-androidqf does;
a zone already known, as when androidqf drives the bugreport inside its
own archive, is kept. Zip entry times are read in that zone and the
mtimes of an unpacked bugreport as UTC instants, so convert_datetime_to_iso
writes both in UTC. Without a zone the wall clock is kept naive and a
warning says so, and an unpacked bugreport gets a warning that its file
timestamps are the extraction's, not the device's. BugReportTimestamps
uses the same reading instead of parsing the properties itself.
A sysdiagnose tarball whose download stopped halfway ends in an EOFError
from gzip while check-sysdiagnose extracts it. Click turns EOFError into
click.Abort, so the command printed nothing but "Aborted!", even with
-v. The extraction now catches the read errors an archive can raise,
names the file and the reason at critical level, says the file may be
truncated or not a gzip tarball, and exits 1. A plain .tar given to the
gzip reader gets the same message rather than a traceback.
* Add a SysdiagnoseInfo module to check-sysdiagnose
check-sysdiagnose had no module of its own: it prepared the archive for
plugin modules and refused to run without one. SysdiagnoseInfo is the
first built-in module. It writes sysdiagnose_info.json with details
about the device and the archive: product type and model, iOS version
and build, serial number, IMEI, MEID and UDID from remotectl_dumpstate.txt
and the mobile activation request, the Apple account name and email from
the App Store daemon database, and the archive's original file name and
creation time from sysdiagnose.log. The build is checked against the
known iOS versions the way BackupInfo does.
The App Store database is copied out of the archive together with its
-wal and -shm sidecars before it is opened, so rows still in the
write-ahead log are read.
With a built-in module the command's list is never empty, so the "no
custom modules" error and its test go. The module joins
IOS_CHECK_IOCS_MODULES like every other module that writes a results
file.
* Note that newer sysdiagnoses lack the App Store daemon database
* Keep refusing check-sysdiagnose runs without a custom module
* Warn instead of refusing when no forensic sysdiagnose module is loaded
MVTSettings.initialise() constructs the settings once with load_env=False
and writes the result to config.yaml, so that values taken from MVT_*
environment variables are never persisted, then constructs them again
with the environment applied.
Since #716 settings_customise_sources() has added env_settings
unconditionally, making load_env dead code. Any MVT_* variable set for a
single run, including MVT_IOS_BACKUP_PASSWORD, MVT_ANDROID_BACKUP_PASSWORD
and MVT_VT_API_KEY, was written in plaintext to config.yaml and read back
on every later run.
Only add env_settings when load_env is true, and add a regression test.
Fixes#915
dumpsys prints a blank line after every namespace block and after a
change history, and the generation registry and any vendor dumps follow
the last block before the section trailer. A record ran until the next
`_id:` line, heading or trailer, so the last row of the section
absorbed those dumps into whichever field came last.
A blank line now closes the record being read; lines that follow it
without an `_id:` are skipped until the next row or heading. The fixture
carries a generation registry after its last block, modelled on the
AOSP dump, and the last row is pinned to exactly its own fields.
A dumpsys settings row can end in `tag:` and, on some vendor builds, in
`isValuePreservedInRestore:` or a bare `notPreservedInRestore` token.
The parser only peeled `default:` and `defaultSystemSet:` off the end
of a record, so on a row without a default those tokens stayed inside
the value, and on a row with one they landed in `defaultSystemSet`. A
value of `0 tag:null` is not the safe value `0`, so a setting at its
safe value was reported as dangerous; these are the false positives
measured in #912.
The metadata keys are ordered fields like the rest of the record, so
`_split_fields` reads them now, which also replaces the separate
handling of the default. The bare token has no `key:` shape and is
printed last, so it is stripped first and recorded as
`isValuePreservedInRestore: false`.
The bugreport fixture gains the three row shapes from #912: a restore
flag after `defaultSystemSet:`, and a `tag:` or a bare token directly
after the value. Two of them sit on dangerous settings at their safe
value, so the unchanged alert count of one is the false-positive check.
The coverage comment never posted: the per-file table with a link on
every file and missing line range exceeds GitHub's 65536-character
comment limit, so PRs got only a badge. Every matrix job also raced to
post the same comment, and the action was unpinned at @main.
Post from one job only, limit the table to files changed in the PR,
drop per-line links, and pin the action. Write the full coverage table
to the job summary as well, which also works for PRs from forks where
the token is read-only.
Set UV_PYTHON from the matrix. Without it, .python-version pins 3.10
and `uv run` rebuilt the venv with 3.10 after `uv sync --python X`, so
all five matrix jobs were testing on Python 3.10.
The `dumpsys settings` parser matched a single regex per line, which
truncated every value that spans more than one line and, when
`defaultSystemSet:` did not fall on the first line, left the trailing
`default:` metadata inside the value. It also keyed results by setting
name within a namespace, so a name recorded twice kept only the last row
and a row without a `pkg:` field was dropped entirely.
Replace it with a line loop that accumulates one record at a time and
splits the `key:value` fields once the whole record has been read.
Results become a list of records carrying the fields dumpsys prints:
namespace, user, _id, name, value, pkg, default and defaultSystemSet,
plus the per-setting change history. History timestamps are printed
without a year, so they are resolved against the "ending at:" time of
the section and serialized into the timeline. This shows which package
changed a security-relevant setting, and when.
The androidqf settings module shares this artifact, so it now emits the
same record shape.
The package version now comes from the latest v* git tag, so a release
is a tag and nothing else. A scheduled workflow tags main on the first
of the month when something was merged since the last release, creates
the GitHub release with generated notes, and publishes to PyPI and the
container registry. Pushing a v* tag by hand goes through the same path.