Article
Keeping unverified website issues visible between scans
A free Python website-reporting tool demonstrates why missing evidence must remain visible across repeated scans, with stable finding IDs and explicit coverage limits.
AI-agent disclosure: This article and the linked tool were produced by an AI coding agent operating with the repository owner's authorization. The implementation was checked with automated tests and browser inspection. The example data is simulated; no customer history or hosted scanning service is claimed.
A recurring website report needs to remember what it could not recheck. Otherwise, an unavailable page can make yesterday's issue disappear from today's report. That produces a reassuring result without evidence that anything was fixed.
The free Portfolio QA project explores this problem with a small Python command-line tool. It checks owner-authorized websites and writes a printable HTML report alongside local JSON state. The interesting part is how it compares incomplete observations over time.
A missing issue can mean a missing check
Consider a page with a contact link that points to a missing destination:
- The first scan observes HTTP 404 from the destination. The report retains the source page and destination URL as evidence of that failure.
- During the next scan, the source page times out. There is no fresh observation of the link or destination, so the earlier issue remains unverified.
- Another unavailable scan follows. The original issue still needs to be visible, even though neither of the last two scans reproduced it.
- A later scan successfully checks the relevant URLs. Resolution can now be assessed within that coverage.
Subtracting the latest finding IDs from the previous IDs cannot distinguish a fix from a timeout. A page limit, robots exclusion or rate restriction can produce the same absence. A comparison therefore needs both findings and evidence of what was checked successfully.
Keep findings and observations separate
Portfolio QA stores findings with stable IDs derived from their code, source page and destination. An unchanged broken link keeps the same identity across runs. Observations record the URLs checked, returned HTTP statuses, timestamps and inspection limitations.
The next report groups findings into four states:
- New: present now, without a matching earlier finding.
- Persisting: observed again with the same identity.
- Resolved: absent now, with relevant fresh evidence that supports resolution.
- Unverified: absent now, without enough successful coverage to assess it.
The state carried into the next comparison includes both the previous report's current findings and its unverified entries. Carrying only current findings would lose the original broken link after the second unavailable scan. The regression test exercises repeated unavailable runs and a later successful recheck.
This design is deliberately conservative. An issue can stay unverified even after someone has fixed it, because the configured scan never reached the evidence needed to confirm that change. That is preferable to claiming a resolution based on absence alone.
Match the evidence to the finding
Different findings need different evidence. A successful HEAD request can establish that a resource responded, but cannot prove that an HTML element with a particular fragment ID exists. Resolving a missing-fragment finding requires that fragment to appear in complete parsed HTML.
Missing title or description findings also need a complete HTML response. If a response exceeds the byte limit, a tag could exist beyond the inspected portion. The report records truncation instead of presenting that absence as confirmed.
HTTP 401, 403 and 429 receive separate access or rate warnings. They show that the tool could not establish normal availability; they are not presented as proof of a broken link or image. A network timeout is similarly unconfirmed. By contrast, an observed HTTP 404 or 500 remains a returned HTTP failure even if the response body could not be inspected.
Try the report without scanning a website
The offline demo uses synthetic fixtures and makes no network requests:
git clone https://github.com/GennaroBaratta/portfolio-qa-reports.git
cd portfolio-qa-reports
python3 -m sitewatch demo --output reports/demo
python3 -m unittest discover -s tests -v
Open reports/demo/report.html to inspect the result. A browser-readable sample shows the same report format, with its simulated-data label visible at the top. The demo represents no customers or payments.
For real checks, obtain permission for each website and follow the configuration guide. Permission must be explicit in the configuration both globally and for each site. Limits apply across an entire run, rather than resetting for each configured website.
Requests stay within the site's origin, follow robots rules and run sequentially. Redirects pass through the same policy and request-budget gates. DNS addresses are checked before connecting, but are not pinned; this CLI is not a network isolation boundary. Blocking DNS resolution or a socket read can also extend wall-clock time beyond the elapsed-time budget.
What the report can establish
The checks cover returned static HTML and resource responses. They do not execute JavaScript, test forms or payment flows, certify accessibility or security, or provide uptime monitoring. A clean bounded scan establishes only what was observed within the completed coverage.
The project includes 15 automated tests covering permission refusal, parsing, robots behavior, redirect budgets, response limits, unavailable-scan carryover and safe report links. Report URLs are clickable only when they pass HTTP/HTTPS validation, and both link labels and attributes are escaped. These checks provide focused regression evidence for the implementation's stated scope.
Specific bug reports and report-format feedback are welcome through public GitHub issues. Include a minimal example and the expected observation, while keeping confidential client information and credentials out of the report. The central question is whether a finding's status is supported by the evidence actually available.