Alterlab Dergipark: Install, Source and Security | FunnelSlayer

Alterlab Dergipark

Published by alterlab-ieu in alterlab-academic-skills

Review recommended55 installs

What this skill does

Harvests article metadata, abstracts, and full-text PDFs from DergiPark (TÜBİTAK ULAKBİM's national journal-hosting platform, 3,000+ journals) via its verified platform-wide OAI-PMH endpoint (https://dergipark.org.tr/api/public/oai/; verbs Identify/ListSets/ListRecords/GetRecord; prefixes oai_dc/oai_mods/oai_marc/oai_etdms; setSpec=journal-slug), parses Highwire citation_* and DC.* meta tags on /pub/{slug}/article/{id} pages, pulls PDFs from the citation_pdf_url path, and emits BibTeX/RIS locall

Add Alterlab Dergipark to your agent

Review the source and files first. When you are ready, copy the prompt instruction or use the CLI command supported by your environment.

Install with a prompt

Paste this into a compatible coding agent:

add this skill "alterlab-dergipark" from https://github.com/alterlab-ieu/alterlab-academic-skills

Install with the CLI

Run this command in a controlled environment after reviewing the repository:

npx skills add https://github.com/alterlab-ieu/alterlab-academic-skills --skill alterlab-dergipark

Skill instructions

DergiPark — National Journal Harvester (TÜBİTAK ULAKBİM)

DergiPark (https://dergipark.org.tr) is TÜBİTAK ULAKBİM's national journal-hosting platform — 3,000+ Turkish scholarly journals (per its home page, 2026-09) on OJS-based infrastructure. This skill harvests its metadata, abstracts, and full-text PDFs reproducibly, by going through the one stable machine surface: a platform-wide OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting) endpoint, plus open article pages. It never guesses bibliographic data — every field comes from the live endpoint, and it emits BibTeX/RIS locally from parsed meta tags.

The interactive search page is gated behind a human-verification challenge, so this skill does not scrape search over plain HTTP. OAI set-harvest and direct article-page fetches are the robust paths (see references/oai-cookbook.md).

When to Use This Skill

Use it when the request is about getting data out of DergiPark:

  • harvest / fetch all articles of a Turkish (DergiPark) journal
  • list a journal's archive or find its slug
  • get BibTeX or RIS (an export format) for a DergiPark article
  • download a DergiPark full-text PDF
  • read a journal's aim-and-scope (öz/kapsam — abstract/scope) or its self-declared index list

Canonical Turkish terms you may see: öz (abstract), kapsam (scope), dergi (journal), künye (citation/masthead). Preserve Turkish spelling with diacritics in all user-facing strings (e.g. Mülkiye Dergisi).

Does NOT Trigger — route adjacent requests to the right sibling

The request is really about…Route to
Whether a journal is indexed in TR Dizin (national citation index) / its TR Dizin statusalterlab-trdizin
Finding/searching a graduate thesis (YÖK Ulusal Tez Merkezi)alterlab-yok-tez
An academic's profile / affiliation (YÖK Akademik)alterlab-yok-akademik
University program admission statistics (YÖK Atlas)alterlab-yokatlas
Computing doçentlik (associate-professorship) eligibility pointsalterlab-docentlik-eligibility
Computing akademik teşvik (academic-incentive) scorealterlab-akademik-tesvik
Depositing a manuscript/data in Aperta (TÜBİTAK open archive)alterlab-aperta
Verifying that cited references exist / hallucination auditalterlab-citation-verifier
Turkish APA / TR Dizin writing-style conventionsalterlab-tr-academic-style

How It Works — Three Scripts

All run with uv run python from the skill directory. No API key. Each prefers requests if installed, else falls back to stdlib urllib. Full recipes and gotchas are in references/oai-cookbook.md; the verified endpoint table is in references/endpoints.md.

1. scripts/dergipark_oai.py — OAI harvest (primary path)

# The slug (== OAI setSpec) is in the journal URL: dergipark.org.tr/en/pub/{slug}
# list-journals only sees the first 100 sets — ListSets stopped paging (2026-09-23)
uv run python scripts/dergipark_oai.py list-journals --tsv | rg -i "mülkiye"

# Harvest a whole journal (follows resumptionToken paging automatically)
uv run python scripts/dergipark_oai.py harvest mulkiye --prefix oai_dc --out mulkiye.json

# One article's structured oai_dc record
uv run python scripts/dergipark_oai.py get 10 --prefix oai_dc

Verbs: Identify / ListSets / ListRecords / GetRecord. Prefixes: oai_dc (parsed into clean fields) and oai_mods / oai_marc / oai_etdms (returned as raw record XML for callers that want the richer schema).

2. scripts/article_meta.py — one article → BibTeX / RIS

uv run python scripts/article_meta.py /en/pub/mulkiye/article/10 --format bibtex
uv run python scripts/article_meta.py /en/pub/mulkiye/article/10 --format ris
uv run python scripts/article_meta.py /en/pub/mulkiye/article/10 --format json

Parses the page's citation_* + DC.* meta tags and formats BibTeX/RIS locally — it does not call any on-site export backend (that route is unverified). The JSON output includes a resolved pdf_url.

3. scripts/journal_info.py — aim-and-scope + self-declared indexes

uv run python scripts/journal_info.py mulkiye --lang en

Returns the aim-and-scope text and the /indexes page, the latter stamped with a DISCLAIMER and cross-check pointers to TR Dizin and DOAJ.

Three Rules That Prevent Wrong Answers

  1. Publication date ≠ OAI datestamp. The OAI <datestamp> is the platform re-index time, not the publication date. To slice by publication year, harvest then filter on dc:date locally — do not rely on from/until.
  2. PDF id ≠ article id. Take the PDF link from the page's citation_pdf_url meta tag (e.g. /en/download/article-file/9 for article /10); never build it from the article id.
  3. Hosting ≠ indexing. Being on DergiPark says nothing about quality or TR Dizin coverage. The /indexes page is self-declared and unverified by DergiPark. For an authoritative TR Dizin status verdict, route to alterlab-trdizin. See references/indexing.md.

Graceful Degradation

Network failures surface as a network_unavailable / oai_error JSON object on stderr (exit 1) — the scripts never fabricate a record or a verdict. If the endpoint is unreachable, say so and stop; do not invent metadata from memory.

References

  • references/endpoints.md — every verified endpoint, prefix, and identifier scheme (with what is explicitly UNVERIFIED), observed live 2026-06-06.
  • references/oai-cookbook.md — copy-paste recipes: slug lookup, full harvest, year filtering, PDF download, and the gated-search workaround via Playwright.
  • references/indexing.md — the DergiPark-vs-TR Dizin distinction and how to cross-check real indexing against TR Dizin and DOAJ.

Part of the AlterLab Academic Skills suite.

Files included

  • evals/evals.json
  • references/endpoints.md
  • references/indexing.md
  • references/oai-cookbook.md
  • scripts/article_meta.py
  • scripts/dergipark_oai.py
  • scripts/journal_info.py
  • SKILL.md

More skills