Search-engine optimisation and XML sitemap for the Islam West Africa Collection (IWAC) Omeka S instance (islam.zmo.de).
The site ships almost no SEO metadata out of the box: pages have a <title> but no
description, no Open Graph / Twitter cards, no canonical link, no structured data, no
Google Search Console verification, and there is no sitemap.xml or robots.txt. This
module fills all of that in — manually for static pages, automatically for every
resource page — and adds a sitemap, a robots file, and an optional IndexNow ping.
It is self-contained, with no third-party runtime Composer dependencies or theme edits. Version 1.1 adds a durable IndexNow outbox table. See operations and migration for the server cron and optional Search Console workflow.
IWAC is bilingual. The same collection is published as two Omeka sites —
afrique_ouest(French, the default the host root redirects to) andwestafrica(English). The module resolves the current site at request time, so every signal (site title, locale, self-referential canonical, breadcrumbs) is emitted in the right language, and each page advertises its other-language counterpart viarel="alternate" hreflang. See Bilingual SEO.
| Area | Detail |
|---|---|
| Meta tags | <title>, <meta name="description">, canonical link on every public page. |
| Open Graph / Twitter | og:title/description/image/type/url/site_name/locale + twitter:card/title/description/image/site so shared links render a rich preview. |
| schema.org JSON-LD | Per-resource structured data typed from the resource class (Person, Place, Organization, Event, NewsArticle, PublicationIssue, the scholarly reference types, VideoObject …), plus WebSite on the home page and BreadcrumbList on resource pages. |
| Citation metadata (Zotero) | Highwire Press citation_* + Dublin Core DC.* <meta> tags so the Zotero Connector, Google Scholar, Mendeley and other reference managers capture each item as a properly-typed reference (newspaper article, magazine issue, journal article, book, chapter, thesis, report, blog post …). |
| unAPI (Zotero RDF) | Primary-source items also advertise a /unapi endpoint serving Zotero RDF. Zotero prefers unAPI over the meta tags, so it imports a fuller record — the call number (Cote) from the iwac- identifier, single-field institutional creators, and Sujet + Couverture spatiale as tags. |
| Item-page citation tools | A "How to cite" resource page block — a formatted Chicago / APA / MLA reference (switchable, copy-to-clipboard) plus BibTeX / RIS / CSL-JSON downloads at /cite/{id}/{format} and the Zotero-RDF link for eligible kinds. Placed via the theme's Configure resource pages screen; the theme renders the UI (its common/citation partial) via the iwacCitation view helper, and this module owns the data. Replaces the BulkExport block for single-item exports. |
| og:image | The large thumbnail of the item's primary media (the page scan / cover); falls back to a site-wide default share image. |
| XML sitemap | /sitemap.xml index → /sitemap-pages.xml, /sitemap-item-sets.xml, /sitemap-items-{n}.xml (5,000 items per file by default). Public resources only, with <lastmod>, <changefreq>, <priority>. Cached. |
| robots.txt | /robots.txt disallowing /admin and pointing crawlers at the sitemap. The staging switch uses page-level noindex, which crawlers can fetch. |
| Google Search Console | Paste the verification snippet in the module config; the <meta name="google-site-verification"> tag is injected site-wide. |
| IndexNow ping | Optionally notifies Bing/Yandex when public content changes (durable, throttled batches and retries). |
Static-page SEO is set by hand (a central admin table); resource-page SEO is derived automatically from each item's metadata, with the configurable defaults filling any gaps.
Both the schema.org typing and the citation typing key off the Omeka resource class (the RDF type), not the data-entry template. IWAC's templates are not 1:1 with classes:
- Template 8 ("Newspaper article") historically held both newspaper articles
(class 36
bibo:Article) and Islamic-publication issues (class 60bibo:Issue); the latter now also has its own template 21, but class is still the reliable split. - The bibliographic references share templates across classes (e.g. template 10 covers
bibo:Book,bibo:EditedBookandbibo:Thesis).
So the maps in config/module.config.php are keyed by class id. This matches how the
IWAC → Hugging Face pipeline derives its subsets (also by class). See the iwac-data
skill (omeka-structure.md) for the full class catalogue.
The module folder must be named IwacSeo inside Omeka's modules/ directory (the
folder name has to match the namespace).
Download IwacSeo-<version>.zip from the
latest release and unzip it into
modules/ — it already unpacks to a correctly named folder and contains only the runtime.
Do not use GitHub's "Source code" archives: they unpack to IWAC-SEO-<version>/, which
the namespace will not resolve. Cloning works too, provided you name the target directory:
git clone https://github.com/fmadore/IWAC-SEO.git modules/IwacSeoThen:
- Admin → Modules → IWAC SEO → Install.
- Configure (see below) — at minimum paste your Google Search Console snippet and pick a default share image.
- Visit
/sitemap.xmland/robots.txtto confirm they render.
No web-server changes are needed: nginx's try_files … /index.php already routes the
.xml/.txt endpoints to Omeka because no such static files exist.
There are no third-party dependencies, so composer install is not required — Omeka
autoloads the IwacSeo\ namespace from src/.
Requirements: Omeka S ^4.2.0, PHP >= 8.2.
Modules → IWAC SEO → Configure. Every field is optional; sensible defaults are applied on install.
| Setting | Purpose |
|---|---|
| Google Search Console verification | Paste the whole <meta name="google-site-verification" …> tag (or just the token). Injected on every public page. Then add the property at search.google.com/search-console and submit /sitemap.xml. |
| Bing Webmaster verification | Same, for the msvalidate.01 tag. |
| Default meta description | Used on pages without their own description (home, browse, search). |
| Default social share image | og:image fallback (≈1200×630). Used when a page or an item with no media is shared. |
| Twitter / X @handle | Emitted as twitter:site. |
| Discourage indexing (whole site) | Staging switch: every page becomes noindex,nofollow while robots.txt permits reading that directive. Off in production. |
| Noindex filtered browse | Keeps filter/sort URLs out of the index; clean pagination stays indexable and self-canonical. On by default. |
| Emit schema.org JSON-LD | Toggle structured data. On by default. |
| Emit citation meta tags | Toggle Zotero / Scholar tags. On by default. |
| Serve unAPI (Zotero RDF) | Toggle the /unapi endpoint + discovery tags for primary sources. On by default; needs the meta tags on for the fallback. |
| Serve /sitemap.xml & /robots.txt | Toggle the sitemap/robots endpoints. On by default. |
| Sitemap cache lifetime | Seconds before the cached sitemap is rebuilt. Default 86400 (24h). |
| Ping IndexNow when content changes | Off by default. See the caveat below. |
| IndexNow key | A hex key you choose (openssl rand -hex 16); served at /{key}.txt. |
Admin → SEO → Static pages lists every site page with editable meta title, meta description, share image (picked with Omeka's asset selector) and indexing (default / index / noindex). Blank fields fall back to the page's own title and the site-wide defaults. Resource pages are not listed — their SEO is automatic.
Because IWAC publishes the collection as two Omeka sites (afrique_ouest/fr and
westafrica/en), the screen opens on the default site and a Site selector at the top
switches between them — overrides are stored per site, so the English pages
(home, about, …) are edited on the westafrica site and the French ones on
afrique_ouest.
Admin → SEO shows what is configured, the sitemap/robots URLs with public-resource counts,
a bilingual (hreflang) coverage report, and a Regenerate button that clears the sitemap
cache so it rebuilds on the next request. The coverage report separates the two ways a page can
be wrong about its translations: broken alternate links — a page_pairs row naming a
counterpart that is not a public page, so the alternate 404s — and pages with no pair, which
simply emit none. Broken is listed first because it is the worse of the two. It also warns when the stored IndexNow key cannot match the /{key}.txt
route (non-hex) and would fail verification.
The head signals are written into Omeka's request-global head placeholder helpers
(headTitle, headMeta, headLink, headScript), which the theme already echoes in
<head> — so no theme template needs to change.
- Resource pages (
view.show.afteron the Item/ItemSet/Media controllers): title fromdisplayTitle(), description frombibo:shortDescription(the AI summary on newspaper articles) /dcterms:abstract/dcterms:description(truncated to ~160 chars), canonical from the resource's site URL,og:imagefrom the primary media, and JSON-LD from the resource class. - Static pages (
view.show.afteron the Page controller): the editor's per-page overrides, else defaults;WebSiteJSON-LD on the home page. - Browse/search pages (
view.browse.after): self-referential canonical and optionalnoindexon filtered variants; tracking parameters are removed and pagination remains indexable. - Every page (
view.layout): site-wide constants (og:site_name,og:locale,twitter:card, verification tags) and gap-fills for anything not already set. Resource values always win because the resource listeners run before the layout listener.
The two canonicals differ deliberately. On the browse routes this module owns, page 2 of a
listing is a real page in a series, so its canonical is self-referential and noindex
is what keeps it out of the index. On a route no phase-1 listener claimed — the IwacSearch
/search app, any other module's controller — the layout gap-fill has no such knowledge, so
it canonicalises to the query-less URL and marks any query-carrying variant noindex, follow. og:url mirrors whichever canonical was written, so a share never disagrees with it.
Turn off Omeka's own JSON-LD embed. Independently of this module, Omeka S core echoes
AbstractResourceRepresentation::embeddedJsonLd()— the resource's entire API representation — into the body of every resource show and browse page. On IWAC that is 157 KB per item (it includes the OCR text) and 2.2–3.9 MB on an authority record, whose@reverselists all ~13,000 items linking to it; Google truncates it and reports unparsable structured data. Disable it per site under Admin → Sites → … → Settings → General → "Disable JSON-LD embed". The data stays available at/api/items/{id}, and Zotero import here uses unAPI and the citation meta tags above, not the embed.
Driven by config/module.config.php → iwac_seo.structured_data.class_types (overridable via
config/local.config.php), keyed by Omeka resource class id:
| Resource class | @type |
|---|---|
foaf:Person (94) |
Person (+ givenName/familyName, affiliation) |
dcterms:Location (9) |
Place (+ geo from curation:coordinates) |
foaf:Organization (96) |
Organization (+ parentOrganization) |
bibo:Event (54) |
DefinedTerm ‡ |
fabio:AuthorityFile (244) |
DefinedTerm (subjects / authority files) |
bibo:Article (36) |
NewsArticle (newspaper article) |
bibo:Issue (60) |
PublicationIssue (Islamic-publication issue) |
bibo:AudioVisualDocument (38) |
VideoObject (+ duration, uploadDate, thumbnailUrl, embedUrl/contentUrl) |
bibo:Document (49) |
DigitalDocument |
bibo:AcademicArticle (35) |
ScholarlyArticle |
fabio:BookReview (178) |
ScholarlyArticle (+ about = the reviewed book) † |
bibo:Chapter (43) |
Chapter |
bibo:Book (40) / bibo:EditedBook (52) |
Book |
bibo:Thesis (88) |
Thesis |
bibo:Report (82) |
Report |
bibo:PersonalCommunication (77) |
CreativeWork (+ isPartOf = the event's URL, when it is a record) |
fabio:BlogPost (305) |
BlogPosting |
† Book reviews are not typed Review. Google's review snippet requires
reviewRating.ratingValue, and an academic book review awards no score, so a Review node
is reported invalid in perpetuity — an error bought for a feature that can never appear. As a
scholarly article about a book, the relationship survives in a form that validates; the
Zotero/Highwire side is unaffected and still declares the review citation kind.
‡ Events are not typed Event, for the same reason book reviews are not typed Review.
Google's Event feature is for events "bookable to the general public" and asks for offers,
performer, organizer, eventStatus and a PostalAddress. All 243 of IWAC's event records
are Notice d'autorité — historical congresses and conferences, none of them attendable — so
they are ineligible by that guideline however complete their metadata, and the type buys 97
errors and some 500 warnings for a rich result that can never appear. DefinedTerm is what
the archive calls them and what class 244 already uses; the 115 Wikidata sameAs links that
do the entity-resolution work are unaffected.
For the same reason a conference paper's isPartOf no longer builds an Event node. That
node was out of range twice over: isPartOf ranges over CreativeWork or URL, and the
event was ineligible anyway. Six of the nineteen papers name a linked event record and keep
the cross-link as its URL; the ten holding only a literal title emit nothing, a name Google
cannot dereference not being worth a range violation.
The Event branch itself is kept in StructuredData because class_types is overridable in
config/local.config.php — an installation whose events are bookable wants that shape, and
it handles the interval case: a dcterms:date may be an ISO 8601 interval
(1979-11-04/1981-01-20, on 55 of the 243 records), which becomes startDate + endDate,
with either side that is not a plain ISO date dropped rather than emitted, since a date a
validator cannot read invalidates the node around it.
thumbnailUrl is taken only from a picture of the resource itself, never from the site's
default share image — unlike image, which may fall back to it. A share card needs some
picture; a claim that the site logo depicts the video is a different kind of statement. Two
sources qualify, in this order: an asset assigned as the item's own thumbnail (the editor's
explicit choice, and the way to give a video whose file yielded no still one that is actually
of it), then the primary media's derivative. A media without derivatives is skipped, because
Omeka answers for it with its generic file-type icon (video.png), which is a picture of
nothing in particular.
A video with no description of its own — no dcterms:description, dcterms:abstract,
bibo:abstract or bibo:shortDescription, which is 310 of the 1,790 — gets one composed
from the record's own facts, in the page's language, since description is required of a
VideoObject and Search Console reports every omission as an error:
Enregistrement vidéo (2 min 49 s) publié par RTB - Radiodiffusion Télévision du Burkina le 15 avril 2022, en français. Lieux : Burkina Faso. Collection Islam Afrique de l'Ouest.
That is a catalogue entry, not prose: who made it (bibo:authorList / dcterms:creator), who
published it, when (to the archive's own precision — a year stays a year), how long it runs,
in what language, where and about what, then the collection. Nothing in it is inferred, which
is the line 1.0.4 drew against summarising the machine transcript. A record that carries any
descriptive text keeps it untouched, and a record with no fact beyond its title gets no
sentence rather than an empty one. VideoDescription holds the wording, in the same EN/FR
string-table style as the citation formatter — the module's services run without a
translator.
A video's uploadDate is emitted only from an explicit dcterms:issued timestamp with a known time zone. Catalogue dates retain their precision in datePublished; missing days, times and zones are never invented. Incomplete video metadata remains an editorial task and may prevent rich-result eligibility.
embedUrl and contentUrl say where the video can be played, and come from disjoint
sources: 1,745 records name a source in fabio:hasURL — all but two a YouTube watch page,
which becomes the /embed/ form — and the 44 digitised from DVD hold the file itself as an
item media, whose original URL is the contentUrl. Only a media Omeka serves as video/*
qualifies: contentUrl must point at content bytes, so a cover image is not a candidate and
neither is a private media. A source URL that resolves to no player (the collection holds one
SoundCloud track and one Wayback capture) yields nothing rather than a link Google cannot
play.
Creative works also carry author/editor/contributor (from bibo:authorList,
dcterms:creator, bibo:editorList, dcterms:contributor), datePublished, inLanguage,
keywords, spatialCoverage, isPartOf (the journal/newspaper/book/event) and publisher
when present. Every record carries sameAs from its dcterms:identifier URI values
(Wikidata, GeoNames, VIAF …) — opaque internal ids like iwac-article-0000001 are skipped.
So that the Zotero Connector and similar tools capture a resource page as a reference,
each item page also emits embedded bibliographic <meta> tags — the two vocabularies
Zotero's Embedded Metadata translator reads:
- Highwire Press (
citation_title, repeatedcitation_author/citation_editor,citation_publication_date,citation_journal_title/citation_inbook_title,citation_volume,citation_issue,citation_firstpage/citation_lastpage,citation_publisher,citation_dissertation_institution,citation_technical_report_institution,citation_doi,citation_language,citation_keywords,citation_abstract,citation_pdf_url,citation_public_url). - Dublin Core (
DC.title,DC.creator,DC.date,DC.publisher,DC.type,DC.language,DC.identifier,DC.subject,DC.description) — emitted for every resource, including authority pages, as a generic fallback.
IWAC's field conventions are baked into the per-kind tags (see src/Service/CitationMeta.php):
- the journal / newspaper / publication title lives in
dcterms:publisher(often a linked item set), notdcterms:isPartOf; - a book chapter's book title lives in
dcterms:alternative; - DOIs live in
bibo:doi(a URI value); ISBN/ISSN are not recorded in IWAC, so those tags are not emitted; - Zotero tags are built from both
dcterms:subject(Sujet) anddcterms:spatial(Couverture spatiale). Both sets are emitted as repeatedDC.subject— the channel Zotero's Embedded Metadata translator turns into tags via its RDF backend (it pre-emptscitation_keywords) — and also folded intocitation_keywordsfor Google Scholar and as a fallback.
The citation kind per class lives in iwac_seo.citation.class_kinds and is overridable.
citation_pdf_url is emitted only for a public PDF media, so restricted bitstreams are
never exposed. Toggle the whole feature with Emit citation meta tags in Configure.
Newspaper articles (class 36) and Islamic-publication issues (class 60) are the bulk of the
archive, but Highwire Press has no container tag for a newspaper or magazine — and any
citation_* container tag forces Zotero's item type to journalArticle (in the Embedded
Metadata translator, the Highwire type wins over DC.type). So those kinds instead:
- set
DC.typeto a valid Zotero item-type id (newspaperArticle/magazineArticle), which the translator's RDF backend honours because no Highwire container tag is present; and - route the publication name through
prism.publicationName→ Zotero'spublicationTitle.
The same DC.type technique types blog posts (blogPost), audiovisual (videoRecording) and
scientific communications (presentation). The scholarly references (journal article, book,
chapter, thesis, report, review) use the standard Highwire container tags, which type them
precisely on their own.
Two things the archive needs simply cannot be expressed through the flat <meta> tags
that Zotero's Embedded Metadata translator reads (verified against translators/RDF.js):
- a call number (French Zotero: Cote) — Zotero only fills
callNumberfrom a typed RDF node (dcterms:LCC), never from a meta tag; a baredc:identifierthat is not an ISBN/ISSN/DOI is dropped; and - a single-field institutional creator — a literal author is always run through Zotero's
cleanAuthor()and split, soAssociation Islamique d'Al Mawadda Burkina Fasobecomes… Burkina / Faso.
So the primary-source kinds (newspaper article, periodical issue, document, audiovisual,
photograph) additionally expose unAPI:
each page carries <link rel="unapi-server" href="/unapi"> and <abbr class="unapi-id">, and
/unapi?id={url}&format=rdf_zotero serves the item as Zotero RDF. Zotero ranks the unAPI
translator (priority 300) above Embedded Metadata (400), so the Connector imports the RDF
instead of scraping the meta tags — which lets the module set every field exactly:
z:itemType→ the precise item type;dc:identifier → dcterms:URI→ the item URL;- Cote ← the
iwac-identifier, as adc:subject → dcterms:LCC → rdf:valuenode →callNumber; - institutional creators ← a
foaf:Personnode with onlyfoaf:surname→ a single-field creator; persons stay literaldcterms:creatorvalues that Zotero splits into first/last; - tags ←
dcterms:subject(Sujet) +dcterms:spatial(Couverture spatiale) asdc:subject; prism:publicationName,prism:number/prism:volume,bib:pages,dc:date,dc:language,dcterms:abstract,dc:rights, and the public PDF as aneprints:document_urlattachment.
The Highwire / DC <meta> tags stay on every page — they still feed Google Scholar and are the
fallback if unAPI is unreachable, and they remain the sole path for the bibliographic-reference
kinds (which Highwire already types precisely). ZoteroRdf reuses the same
iwac_seo.citation.class_kinds map as CitationMeta; toggle the endpoint with Serve unAPI
in Configure (the iwac_seo_unapi setting).
The capture above is for reference-manager software. Human readers get a public "How to cite"
panel, provided as a resource page block: this module registers it (resource_page_block_layouts),
so it appears in the theme's Configure resource pages screen (Admin → Themes → your theme →
Configure resource pages) and an admin controls its region and order. The block delegates to the
theme (IWAC-theme) for the UI (its common/citation
partial + styling + JS) while this module owns the citation data, so the resource-class → kind
mapping is never duplicated. The block reads its data through the iwacCitation($item) view
helper, which returns:
- a formatted Chicago (default), APA and MLA reference — switchable, one-click copy —
produced by
CitationFormatter. It is hand-rolled (no CSL-processor dependency; the module ships no bundledvendor/) and bilingual: connectives and month names follow the site language, so the French site reads Dans, sous la dir. de, 7 décembre 2018; - downloads at
/cite/{item-id}/{format}— BibTeX (.bib), RIS (.ris) and CSL-JSON (.json), served byCitationControlleras anattachmentwhose filename is theiwac-accession id; and - the Zotero RDF link (the
/unapiendpoint above) for the Connector-eligible kinds.
This is the single-item replacement for Daniel-KM/Omeka-s-module-BulkExport. All three
formatters and serialisers read one normalized record from CitationData, which reuses the
CitationMeta field conventions (container in dcterms:publisher, a chapter's book title in
dcterms:alternative, NumericDataTypes dates, person-vs-institution creators). Authority records
(person, place, organisation, event, subject) are not citable works — the helper returns
null and the panel is hidden. The panel and downloads share the citation kill-switch with the
meta tags (the iwac_seo_citation_meta setting).
IWAC is the same collection under two Omeka sites — afrique_ouest (fr) and westafrica
(en). The module ties them together the way Google expects:
- Canonicals are self-referential per language. The French page canonicals to its own
…/afrique_ouest/…URL, the English page to…/westafrica/…. It never points one language at the other (that would drop a language from the index). rel="alternate" hreflanglinks every page to its counterpart — with a self link and anx-default(the French site, matching the host-root redirect) — so Google serves the right language and treats the pair as one set instead of duplicate content.-
Resources (items, item sets, media) are shared across both sites under the same
o:id, so the alternate is just the same path under the other site slug — fully automatic. -
Static pages have different slugs per language (
accueil/home,a-propos/about…); they are mapped by theiwac_seo.hreflang.page_pairstable inconfig/instance.config.php. A page with no entry simply gets no alternate (never a broken one). Update the table when pages are added or renamed.The table is generated, not authored. The Internationalisation module records each page's counterpart and exposes it on the REST API as
o-module-internationalisation:related_page, so that module is the only place a pairing is written down; the committed table is a build product, which is what keeps a page render free of any lookup.composer hreflang:fixregenerates it, andcomposer hreflang:checkaudits it — reporting pairs the module records but the table omits, rows that contradict the module (which emit a 404 alternate — worse than emitting none), and rows naming pages that no longer exist. Rows are sorted and aligned deterministically, so regenerating an already-correct table rewrites nothing.Adding or renaming a page is what invalidates the table, and nothing in this repository changes when that happens — so the
hreflang driftworkflow regenerates it weekly and opens a pull request if anything moved (a pull request, not a push: the table drives what search engines are told about every page). The same workflow audits, without rewriting, any PR that touches the table. It needs the network, so it is deliberately outsidecomposer check.
-
og:locale+og:locale:alternateadvertise both languages to social platforms.- The item / item-set sitemaps carry
<xhtml:link rel="alternate" hreflang>for each language, so both versions are discoverable from the single (default-site)<loc>entries.
Configure the language map, x_default and page pairs under iwac_seo.hreflang (override via
config/local.config.php); set enabled => false to turn it all off.
/sitemap.xml— sitemap index listing the children below./sitemap-pages.xml— the home page + all public site pages (driven by the site navigation, so menu order and depth set the priority)./sitemap-item-sets.xml— public item-set browse pages./sitemap-items-{n}.xml— public items, defaulting to 5,000 URLs per file./sitemap-pages-{site-slug}.xml— other public language sites, including untranslated pages; pages markednoindexare excluded.
Each item entry also carries an <image:image> element (the primary media's large
thumbnail — the page scan or cover) so Google Images can index the scans; disable via
iwac_seo.sitemap.include_images in a local config override.
Resource ids + modified timestamps are read with one lean DBAL query per type (public
resources scoped to the site), avoiding representation hydration for each item. Output is
cached under files/iwac-seo-cache/ and served with Cache-Control / Last-Modified
headers; the cache is invalidated when an item or page changes, and any cache failure falls
back to live generation.
Public content changes enter a durable outbox. Background jobs send leased batches of up to 200 URLs, grouped by origin, and acknowledge only accepted submissions. Failed work retries with backoff; leases expire after worker failure. Schedule the independent five-minute drain so pending work does not depend on another edit. See operations for migration, cron, retries and the optional GitHub Search Console monitor.
Google is not pinged: its sitemap-ping endpoint was retired in 2023. Google discovers
content via the robots.txt Sitemap: line and Search Console.
IwacSeo/
├── Module.php # listeners, ACL, install/upgrade/uninstall, config form
├── composer.json # type omeka-module; PSR-4 IwacSeo\ -> src/ (no runtime deps)
├── config/
│ ├── module.ini
│ ├── module.config.php # wiring: routes, services, navigation
│ └── instance.config.php # IWAC data: class maps + page translations
├── src/
│ ├── Controller/
│ │ ├── SitemapController.php # /sitemap*.xml, /robots.txt, /{key}.txt
│ │ ├── UnapiController.php # /unapi -> Zotero RDF
│ │ ├── CitationController.php # /cite/{id}/{format} -> BibTeX/RIS/CSL-JSON
│ │ ├── Admin/SeoController.php # dashboard + static-page table
│ │ └── Concern/SendsResponses.php # shared XML/text/file response plumbing
│ ├── Form/{ConfigForm,PageSeoForm}.php
│ ├── Job/PingSearchEngines.php
│ ├── View/Helper/Citation.php # iwacCitation() -> "How to cite" view-model
│ ├── Site/ResourcePageBlockLayout/Citation.php # "How to cite" resource page block
│ └── Service/
│ ├── HeadMetadata.php # decides every <head> SEO signal
│ ├── HeadWriter.php # writes them; tracks what the request has set
│ ├── StructuredData.php # schema.org JSON-LD (by resource class)
│ ├── CitationKind.php # the kind vocabulary + export type tables
│ ├── CitationKindMap.php # resource class id -> kind (one shared map)
│ ├── CitationMeta.php # Highwire + Dublin Core citation tags
│ ├── ZoteroRdf.php # Zotero RDF (served via unAPI)
│ ├── CitationData.php # builds a CitationRecord from an item
│ ├── Citation/ # CitationRecord, Creator, IssuedDate (the record)
│ ├── CitationFormatter.php # Chicago / APA / MLA text (hand-rolled, bilingual)
│ ├── CitationExport.php # BibTeX / RIS / CSL-JSON serialisers
│ ├── Concern/ResourceValueReader.php # shared value-readers (trait)
│ ├── SitemapGenerator.php # which URLs go in which sitemap
│ ├── Sitemap/ # SitemapRepository, UrlsetWriter, XmlCache, SitemapDocument
│ ├── PageSeoStore.php # per-page overrides (site setting)
│ ├── PingQueue.php # IndexNow queue: dedupe, flood cap, throttle
│ ├── Pinger.php # IndexNow submit
│ ├── SettingsGate.php # typed reads over the iwac_seo_* settings
│ ├── ResourceUrl.php, ViewLocale.php, Text.php # small shared helpers
│ └── *Factory.php
├── view/iwac-seo/admin/seo/{dashboard,pages}.phtml
├── asset/css/admin.css
├── language/ # template.pot + fr.po/fr.mo
└── tests/ # unit + real-Omeka integration suites
curl -s https://islam.zmo.de/sitemap.xml | head→ a<sitemapindex>; the child sitemaps list<url>entries with<lastmod>.curl -s https://islam.zmo.de/robots.txt→Disallow: /admin/and aSitemap:line.- View-source an item page (
/s/afrique_ouest/item/2231): confirm<title>,description,og:*,twitter:*,<link rel="canonical">and anapplication/ld+jsonblock. Validate the JSON-LD with the Rich Results Test. - Open a reference page (e.g. a journal article) with the Zotero Connector and confirm it saves as the right item type with authors, date, container and DOI; open a newspaper article and confirm it saves as a newspaper article with the newspaper as the publication.
- Home page:
og:site_nameand aWebSiteJSON-LD block; the GSC tag once a token is set. - In Search Console: add the property, verify via the meta tag, submit
/sitemap.xml.
- COinS (
<span class="Z3988">) as a complementary reference-embedding signal, and to expose references on list pages. - Per-URL hreflang in the pages sitemap too (static-page alternates are already emitted
on-page; only the item / item-set sitemaps carry
<xhtml:link>so far). - Optional nginx-level caching of
/sitemap*.xml(the module already emitsCache-Control/Last-Modified; the image-sitemap entries shipped in 0.6.0).
Removes every iwac_seo_* global setting, the per-site iwac_seo_pages overrides and the
sitemap cache directory (files/iwac-seo-cache/).
No production dependencies; composer install pulls the PHPUnit, PHPStan and
PHP_CodeSniffer development tools. Run every local quality gate with:
composer install
composer checkThe suite covers the pure logic that regresses most silently — the Chicago/APA/MLA
formatter, the BibTeX/RIS/CSL-JSON serialisers, hreflang resolution and the text
utilities, plus sitemap policy and IWAC's instance-configuration contracts. GitHub Actions
(.github/workflows/ci.yml) runs syntax checks and PHPUnit on PHP 8.2–8.5, then checks
Composer metadata, translations, PSR-12 and PHPStan at the declared PHP 8.2 floor. A
separate PHP 8.5 job installs Omeka S 4.2.1 and exercises CitationMeta, ZoteroRdf,
HeadMetadata, the service/controller managers and public routes against the real Omeka
and Laminas classes. It intentionally does not add framework dependencies to this module
or expand the narrow unit-test shims.
To run that boundary locally, install Omeka separately and point OMEKA_PATH at its root:
OMEKA_PATH=/path/to/omeka-s composer test:integrationOmeka's vendor/autoload.php is loaded before the module's development vendor, matching
the production ownership of Laminas and PSR classes.
ROADMAP.md documents the refactoring plan behind 0.6.0 and what remains deliberately
out of scope.
GPL-3.0-or-later. © Frédérick Madore.
Use the Cite this repository button on GitHub, or CITATION.cff
directly — it carries the ORCID and the released version, and GitHub renders it as APA or
BibTeX. Keep its version and date-released in step with config/module.ini when
releasing; the release workflow fails the build if they disagree.