Publication List Generator

Paste one snippet into your site and the publication list on it keeps itself up to date. It is a few lines of HTML — into a WordPress block, a lab page, a departmental template, anywhere you can put markup — and from then on the list is rebuilt from ORCID, PubMed and researchmap in each visitor’s own browser, every time the page loads. There is no server, no account, no API key and nothing to renew, and nothing to remember to do again in a year. That last part is the point: a publication page that has to be edited by hand stops being edited, and a lab page whose newest paper is from 2022 makes an active group look dormant.

It also does the smaller job: paste the PMIDs and DOIs cited in an article and get a formatted reference list back, in Vancouver, APA, Harvard, Chicago or Nature style, ready to copy into Word, WordPress, BibTeX or RIS.

How to use it

The wizard opens on three modes. Pick the one that matches what you are making:

  • Reference list — paste the PMIDs and DOIs cited in an article and get a formatted reference list.
  • My publications — an auto-updating list of one person's work, from ORCID, researchmap and PubMed.
  • Lab or group — several members at once, plus pinned individual papers and a review step for anything uncertain.

What you give it are seeds: ORCID iDs, researchmap permalinks, and PubMed queries. In the Lab or group mode each member is a row — Name, ORCID iD, researchmap, and the Joined / Left fields and Freeze button described below — with Add member for the next one. A list that already exists in a spreadsheet goes in through Paste a list from a spreadsheet, collapsed under the rows: it finds the identifiers by shape rather than by column, so a paste works in either order (Name → ORCID or ORCID → Name), a header row is discarded, and importing adds to the rows instead of replacing them.

The pinned papers box for PMIDs and DOIs sits above the PubMed queries, and deliberately so. A pinned paper is confirmed outright — it goes onto the list whatever else the configuration says, which is the difference between naming a paper and searching for one. A query is the harder instrument: it can be too broad, it can return nothing, and what it finds waits in a queue.

The options are shared across the three modes:

  • Citation style: Vancouver (the default), APA, Harvard, Chicago or Nature.
  • Grouping: the default is a section per publication type — Original Articles & Reviews, Letters, Editorials, Other Publication Types — with a year divider inside each. You can also have the type headings alone, the year headings alone, or one flat numbered list, which is what a reference list wants.
  • Heading level: automatic by default. The script measures the page it has been pasted into, finds the nearest heading above the list and renders one level below it, so a list dropped under an <h2> writes its sections as <h3> and fits that page's outline. It is clamped to H2–H5 at both ends, and the year dividers always sit one level under the section headings. Pick a fixed level instead if you would rather decide it yourself — and note that baking the list into the snippet takes automatic away, because a copy written here cannot know the page it will end up on; the level then becomes an explicit H3 in both halves, so the headings do not move when the script refreshes the list.
  • Preprints: off by default. A publication list normally means published work. Tick Include preprints and they come back in a section of their own.
  • Japanese-language records: by default they get a trailing section of their own. They can also be interleaved with everything else, or hidden.
  • Period: a from and a to, either YYYY or YYYY-MM. Both start empty — there is no default period, and leaving them alone means the whole record.
  • Bold names: the names to embolden inside each citation. Spell them out in full — Yuki Furukawa, not Furukawa Y, because a short form cannot tell Yuki from Yuri. The tool warns you when a bold name lands on two different people.

Then press Generate, and read the warnings panel before you read the list. It names which sources failed, which pinned identifiers could not be retrieved, which records were dropped as errata, and which preprints were held back — by title, not just by count, so you can see whether something you expected is on that list rather than on the page.

Confirmed records and candidates

A PubMed name search catches other people. Furukawa Y[au] returns every Furukawa Y in PubMed, which for a common Japanese or Chinese surname is a great many strangers. So the tool sorts what it finds into two piles.

Confirmed is everything from ORCID, from researchmap, from a pinned PMID or DOI, and from a PubMed [auid] search — that last one is a search on an ORCID identifier rather than on a name, so it is trusted for the same reason ORCID is. Candidates are what a name search turned up and nothing else corroborates. They land in a review queue, and you approve or reject each one. Decisions are stored, so the same paper is never put to you twice: approvals become pins and rejections become exclusions.

The consequence worth understanding before you paste anything into your site: a candidate never appears in an embed. The review queue exists in the wizard and nowhere else, because there is nobody at the far end of a page load to work through one. Under the default policy an undecided candidate is absent from your page permanently, not until the next refresh. The wizard tells you the count beside the snippet, and if every record is an undecided candidate it declines to generate a snippet at all rather than hand over markup that would render an empty list for ever.

There is a setting that publishes name-search hits without review. It is off by default and it should stay off for a name search; it is how another Tanaka H's cardiology paper ends up in your neuroscience group's list. The narrower version of the same idea is safer: a single query can be marked trusted, once, after you have run it on PubMed and read every result — that is an assertion you make about one query, not a policy for all of them. The tick travels in the snippet beside the query rather than inside it, as the positions of the ticked queries, so trusting one changes nothing about how you install the list.

When someone joins or leaves

This is the problem every lab page eventually has, and it is one-sided. Adding a member is trivial. Removing one is not, because their ORCID record follows them: leave the seed in place and it goes on finding everything that person publishes, including the papers they write at their next institution, which then appear on your page as your group's work. Deleting their line the day they leave is the obvious fix and it is also wrong — work done in your group is routinely published a year or two after the person has gone, and that work is yours.

In the Lab or group mode every member's row carries Joined and Left date fields and a Freeze button, beside the name and the identifiers. They are on screen before you have typed anything, so you can see the capability exists rather than discovering it eighteen months later.

  • Freeze is the reliable answer. Generate the list, then press it: their publications so far become explicit pins — by DOI or PMID, exactly as if you had typed the identifiers — and their seed comes out of the configuration. Nothing is inferred, so nothing can be inferred wrongly. The row tells you how many papers it will pin before you confirm, and names any that have neither a DOI nor a PMID, since those are the ones that will drop off. It is recoverable: freezing comments the member out rather than deleting them, and the commented line is still there under Edit the member list as text, so deleting the # puts the seed back.
  • Joined / Left dates are the fallback for when nobody remembers to freeze anybody. A departed member stops contributing new work after a grace period, which defaults to 24 months — an estimate of ordinary publication lag through submission, review, revision and production, not a rule derived from anything. Leave Left blank for anyone still in the group; a member with no dates at all has no time limit, which is what every member is until you touch these fields. A different grace period is the one thing the row cannot say — write it in the text box underneath as 2019-04..2023-03+36 on that member's line.
  • Remove sits on each publication's own line. Freezing pins whatever is on the list at that moment, so it can pin a paper from the member's new institution that was already showing; Remove takes it off, and an exclusion outranks a pin whether the pin was typed by you or written by the freeze. What you removed is listed above the list with an Undo beside it.

Dates are a rule about dates, so they can be wrong about an individual paper in both directions. Two things they never do: a pinned paper is never removed by a window, and a paper co-authored by a departed member and a current one keeps the current member's claim and stays — the filtering is applied per member, not per paper. Everything a window removes is named in the warnings panel with the window responsible.

What you can take out

Six things to copy and two to download:

  • Copy All (for Word) — the formatted list on the clipboard with its styling intact, for pasting into a document. It arrives with a note at the top addressed to you rather than to your readers, asking you to check the list over.
  • WordPress blocks — block markup, so the list stays editable in the WordPress admin screen.
  • Static HTML (no auto-update) — the finished list as plain markup. No script tag, no wrapper, no attributes: nothing for a sanitiser to strip and nothing for a browser to execute.
  • Markdown, BibTeX and RIS.
  • Downloads of publications.bib and publications.ris, for a reference manager that would rather be given a file than a clipboard.

Putting the list on your own site

Three routes, none of which needs an account, an approval, a hosted file, or a pull request to anybody's repository. Everything a list needs travels in the snippet you paste.

RouteUpdates itselfIn your page's HTMLNeeds a script tag
Script snippetyesoptional snapshot, then liveyes
iframeyesnono
Static HTMLnoyesno

The script snippet is the recommended one, and pasting it into a Custom HTML block is the whole installation. By default it is short: a <div> carrying the settings as attributes, two small-print lines, and a <script> tag. The script fills the container in on load.

There is a tick box — Include the list itself in the snippet — that also bakes the list as it stands into the markup. It is recommended and it is off, which is a deliberate pair rather than a contradiction: it is worth having, and it is also most of the snippet's length, and a wall of markup is what stops somebody pasting it at all. What ticking it buys is that the list is in your page's own HTML, so search engines index it, visitors with JavaScript disabled read it, and it is on screen at the first paint rather than when the fetch lands. The two small-print lines — where the list came from, and the credit — sit outside that baked copy either way, so the short snippet still carries both.

The other deliberate thing: the script never blanks the list. If ORCID is down, if the network fails, if the script never loads at all, whatever is in the container stays on the page. Every class is namespaced publist- and the markup is unstyled, so it inherits your site's typography and you can style it from your own stylesheet.

The iframe snippet is the fallback for a CMS that strips script tags, which is common in university systems. It works the same way and updates the same way, but its content is a separate document, so it is not in your page's HTML: search engines will not index it as part of your publications page, and a visitor with JavaScript disabled sees an empty frame. The snippet carries a small inline listener that resizes the frame to its content; if that gets stripped too the frame still works, at its fixed fallback height.

Static HTML is for a page that should not run anything at all — an intranet with a strict content policy, a departmental template whose only editable thing is a rich-text field. It makes no external requests when the page renders. What you give up is in the name: it is a snapshot of today, and you regenerate and re-paste when you have new papers.

There is no size or shape of configuration that the snippet cannot carry, so there is no hosted-file route and nothing to keep in sync. A lab with twenty members, a long exclusion list, a query you have marked trusted: all of it is in the attributes. A comma inside a value — which a realistic PubMed query has, as in Furukawa Y[au] AND (Tokyo, Japan[ad]) — is written %2C on the way out and read back as a comma on the way in, so the query survives a comma-separated attribute intact and comes back into the wizard exactly as you typed it.

Coming back to a list you already made

You do not have to rebuild a configuration by hand a year later, and there is nothing to save except the snippet itself: the snippet is the configuration, and a copy of it in a text file, in an email to yourself, or on the page it is already living on is a complete backup. Paste it back into the wizard — "Start from an existing snippet", above the mode tabs — and every setting it carries is read back out and put into the form. The whole snippet works, the opening <div> on its own works, and so does an iframe snippet. A URL is not accepted, and nothing is fetched from one. Nothing is built until you press Generate, so you see what came back before anything happens, and if something could not be recovered the wizard names it rather than quietly returning a different list.

A live example

Below is the output, embedded: one ORCID iD and one researchmap permalink, Vancouver style, grouped by publication type with a year divider inside each — the defaults, with nothing turned on for the occasion. It is being rebuilt in your browser as you read this page, which is also why the count is whatever the sources hold today rather than a number I typed.

Method and sources

Five public APIs are called, all of them directly from the visitor's browser. Three of them are seeds — they decide which works belong on the list — and two are enrichment, which fill in metadata for works the seeds have already claimed.

  • ORCID [1] — the primary seed, and the researcher's own curated works record.
  • PubMed / NCBI E-utilities [2] — a seed by ORCID identifier ([auid]) or by query; also supplies PMIDs, journals, dates, languages and publication types.
  • researchmap [3] — a seed for Japanese researchers, and the only source here that reliably carries Japanese-language journal articles.
  • OpenAlex [4] — enrichment only, keyed on the DOIs, PMIDs and titles the seeds produced. It supplies author names, work types and missing metadata. It is deliberately not used as a seed: its author disambiguation is not accurate enough to decide whose paper a record is.
  • Crossref [5] — peer-review status for F1000-family open-review journals, where an article can be posted before review is finished, and full author names when OpenAlex has none.

Records describing the same work are merged rather than discarded, by DOI, then by PMID, then by normalised title and year; versioned DOIs of the same work collapse into one entry. The consequence of all of this is worth stating plainly: the list is a reformatting of what has been registered. It is not a literature search, and it will not find a paper that nobody registered anywhere.

What the page contacts, and what it does not

Everything runs in the visitor's browser. There is no backend to this tool: two static files are hosted, the script and the wizard, and nothing else. That is also why the honest version of "it all runs in the browser" is longer than one sentence — running in the browser is precisely what makes the visitor, rather than a server of mine, the party talking to the upstream APIs.

What a page carrying the snippet actually requests is the script itself from GitHub Pages, and then, depending on which seeds you configured, pub.orcid.org, eutils.ncbi.nlm.nih.gov and api.researchmap.jp, plus api.openalex.org for author names and work types and api.crossref.org for peer-review status. There is no proxy in front of any of them. Each receives what any third-party resource on a web page receives: the visitor's IP address and User-Agent, and, under browsers' default referrer policy, your page's origin rather than its full URL. It is the same exposure as an embedded font or a hotlinked image. What they do not receive:

  • No cookies. The embed sets none and sends none.
  • No identifier of the visitor. The only identifiers in any request are the ORCID iDs, permalinks, queries, DOIs and PMIDs you configured — which identify the researchers whose list this is, are already public, and are the point of the list. Nothing identifies the person reading the page. No contact email and no API key is sent to any of the five.
  • Nothing reaches me. No backend, no analytics, no telemetry, no error reporting: not a beacon, not a pixel, not a logging endpoint. One caveat in the other direction: the script is a static file on GitHub Pages, so GitHub serves it and sees that request the way any CDN does. The file is self-contained, so hosting your own copy removes even that.

No account, no sign-up, no API key. Built lists are cached in the visitor's own localStorage for 24 hours, under a namespace of the tool's own; a storage failure degrades to "no cache" rather than to a broken page. If your page must not talk to third parties at all, the Static HTML output makes no external requests whatsoever — that is the reason the button exists.

Limitations

This is a research tool, and the output always needs a human read before it is published. The failure modes are known:

  • It can only show what is registered. A paper missing from ORCID, researchmap and PubMed alike is missing from the list, and errors in those records are reproduced faithfully. Gaps in an ORCID record become gaps on the page, and the fix is in ORCID rather than in any workaround here.
  • PubMed name searches catch other people. That is what the confirmed/candidate split and the review queue are for, and why candidates are hidden by default. A query that returns PubMed's 200-result cap is flagged as probably too broad; narrow it with an affiliation ([ad]) or a date range ([dp]).
  • A group name searched in [au] comes up empty, and that proves nothing. PubMed files a collective author — a study group, a trial consortium — in a field of its own. Measured against the live API on 6 August 2026: "RECOVERY Collaborative Group"[au] returns 0 records, and "RECOVERY Collaborative Group"[cn] returns 18, which PubMed translates as Author – Corporate. An [au] search finding nothing is not evidence that a group is absent from PubMed; check [cn] before concluding anything. If [cn] is empty too, the journals never supplied a collective name for those articles and no rewording will reach them — pin the papers by PMID or DOI instead. [cn] is still a name rather than an identifier, so it is not auto-trusted the way [auid] is.
  • A candidate never appears in an embed. The wizard preview is therefore a superset of what your page will show, and the wizard says so with the count.
  • Type classification is imperfect. The categories come mostly from OpenAlex work types, falling back to what ORCID or researchmap reported, and there is no cross-source vote to catch a mistake. Preprint servers are recognised by journal name against a fixed list, so a server not on it will be miscategorised. Setting the grouping to year, or to none, sidesteps categorisation entirely.
  • An F1000-family article that Crossref does not yet report as approved is filed as a preprint, and therefore hidden by default along with the rest of them. That is the intended reading — it has been posted, not yet peer-reviewed — but it is a surprise if you were not expecting it, which is why every held-back record is named in the warnings.
  • Author names are only as good as the source. ORCID work summaries carry no author list at all, so author names come from OpenAlex enrichment; researchmap stores short forms in a field that reads like a full-name field, and its author ordering varies between accounts.
  • Group membership is not something the sources know. Neither ORCID nor PubMed will tell you that a student left in 2023, so nothing here can work it out on its own. Freezing is explicit and reliable; dates are a rule and can be wrong about an individual paper.
  • Dates recorded for the same work differ between sources, so a paper published near a year boundary is worth checking against the original record before an annual list is finalised.

How to cite

This is a formatting utility rather than an analytic method, so in most uses it needs no citation. Where the provenance of a list has to be documented — an annual report, a grant application, an institutional review — a sentence such as the following is enough:

The publication list was generated from the members' ORCID, PubMed and researchmap records, with metadata completed from OpenAlex and Crossref, using the Publication List Generator (Furukawa Y, https://yukifurukawa.jp/publication-list-generator/), accessed [date], and checked manually against the original records.

The tool is open source under the MIT licence.

References

  1. ORCID. https://orcid.org/
  2. National Center for Biotechnology Information. Entrez Programming Utilities Help. https://www.ncbi.nlm.nih.gov/books/NBK25501/
  3. researchmap. https://researchmap.jp/
  4. OpenAlex. https://openalex.org/
  5. Crossref. https://www.crossref.org/