What a web directory is
A web directory lists websites in a hierarchy of topics, each listing chosen and described by a person rather than ranked by an algorithm. The early web relied on directories to find anything; their categories (Arts, Business, Science and so on, down to narrow subtopics) were one of the first large classifications of the web.
The Open Directory Project
The project began in June 1998 as GnuHoo, was soon renamed NewHoo, and a few months later was acquired by Netscape, which renamed it the Open Directory Project. It later ran under the domain dmoz.org (from "directory.mozilla.org"), which became its common name. When Netscape became part of AOL, AOL continued to host it.
- Volunteer editors: tens of thousands of editors over its lifetime, each responsible for categories they knew, reviewing submitted sites and writing short neutral descriptions.
- Scale: millions of listed sites in a deep topic tree, with separate trees for many languages.
- Open data: the whole directory could be downloaded under a permissive license (the Open Directory License, later also offered under Creative Commons Attribution), so other sites could reuse it.
Its open license is what made it influential. Search engines and portals republished it; Google built its Google Directory on ODP data until it closed that service in July 2011. For years an ODP listing was a sign that a human editor had looked at a site.
The data format
The ODP published two dumps: structure.rdf.u8 with the category tree and content.rdf.u8 with the listed sites, their titles and descriptions. The format was RDF-like XML designed before the RDF standard was final, so it needed its own parsers. It remains a well-known dataset for research on web classification.
<!-- A simplified entry from the ODP content dump (content.rdf.u8) -->
<Topic r:id="Top/Science/Software">
<link r:resource="https://www.example.org/"/>
</Topic>
<ExternalPage about="https://www.example.org/">
<d:Title>Example Project</d:Title>
<d:Description>An example description written by a volunteer editor.</d:Description>
<topic>Top/Science/Software</topic>
</ExternalPage>Closure, and Curlie
AOL closed DMOZ on March 17, 2017. Former editors kept the directory alive and relaunched it under a new name, Curlie (curlie.org), which continues as a human-edited, multilingual directory run by volunteers.
What directories meant for structured knowledge
- Human curation at scale. The ODP showed that a large community could maintain a shared classification, the same model Wikipedia and Wikidata later used for knowledge.
- Open licensing. Releasing the data openly let it spread into many products, as CC0 and CC BY-SA data does today.
- Categories, not entities. A directory classifies pages under topics; it does not say what a page is about in machine-readable terms. Structured data such as schema.org and knowledge bases such as Wikidata took that next step: identifying the entities themselves, including people, and the relations between them.
SelfBadge is not a directory in the ODP sense: it does not list websites or rank them. Its people directory lists profiles, each with structured data about one person.
Specifications and sources
Related reference
- Open data sources about people: Wikidata, DBpedia, OpenAlex and ORCID public data, and their licenses.
- Entity resolution: how knowledge graphs merge records about the same person.
- How AI assistants identify people: the signals assistants use, and why namesakes get mixed up.
- All reference pages