What I actually want to know
How smart can a model get on data that specifically opted in? A generalist wants everything. A specialist wants one domain done properly, and code or markdown may be narrow enough that strict permission still leaves a usable amount of it.
Answering that needs a corpus filtered on permission alone, at a size worth measuring, and no such corpus of third-party text exists. Building the thing that could collect one is what the rest of this page is about. RQ5 is where the measurement would land.
The hypothesis
My hypothesis is that access, use, risk, provenance and revocation are five separate decisions, each backed by its own evidence. The alternative is that the signals publishers actually ship are too fragmented to decide anything automatically.
If the alternative holds, this stays an observatory and an allowlisted collector.
The questions
- RQ1
Signal prevalence
How often is each mechanism present and valid?
- RQ2
Conflict and ambiguity
How often do mechanisms disagree?
- RQ3
Adapter behaviour on bad input
Can each adapter survive malformed input?
A conformance corpus and a differential test against an independent implementation. That exists for robots only.
- RQ4
Politeness
How much load does one origin actually take?
- RQ5
What survives the filter
What fraction survives rights and risk filtering?
Unmeasured on any site the operator does not own. This is the number the question above turns on.
- RQ6
Privacy, secrets and safety
How many allowed documents carry secrets or personal data?
- RQ7
Revocation and provenance
Does a deletion drill run end to end?
A drill from adding the exclusion through to propagating it downstream. The downstream steps are unimplemented.
Six workstreams
- Workstream A
The policy observatory.
Measuring what publishers declare.
Largely delivered. - Workstream B
Legal review.
Decision records and the open questions.
Partial. - Workstream C
The allowlisted pilot.
Collection under scoped grants only.
Partial. - Workstream D
Data quality and deduplication.
Separate components, no pipeline yet.
Partial. - Workstream E
Scale and reliability.
Unstarted. - Workstream F
Dataset registry and training handoff.
Unstarted.
Legal scope
The crawler stores evidence. Legal analysis is a separate job, and I do not interpret anyone's terms of service. The thresholds that stop the project are on the charter.