Reference Extractor
Upload a Word (.docx) or text file — or paste your paper — to pull out the reference list, find every in-text citation, collect DOIs, and spot sources that are cited but missing from the list. Free, no account, and your document never leaves your browser.
How the extractor works
- Your .docx is unzipped and parsed locally with built-in browser APIs — nothing is uploaded to any server.
- The tool looks for a References, Bibliography, or Works Cited heading and splits everything below it into entries (wrapped hanging-indent lines are rejoined automatically).
- The body text is scanned for author–date citations like (Smith, 2020) and Smith (2020), and every DOI in the document is collected.
- In-text citations and reference entries are cross-checked both ways, so orphaned citations and never-cited entries stand out before you submit.
Frequently Asked Questions
- What does the reference extractor do?
- It reads your paper — a Word (.docx) file, a .txt file, or pasted text — finds the References / Bibliography / Works Cited section, and splits it into individual entries. It also scans the body for author–date in-text citations like (Smith, 2020), collects every DOI, and cross-checks the two lists so you can see citations that are missing from the reference list and references that are never cited.
- Is my document uploaded to a server?
- No. The extractor runs entirely in your browser — the .docx file is unzipped and parsed locally with built-in browser APIs, and nothing leaves your device. It is free, with no account and no file-size paywall.
- Which file types are supported?
- Word .docx files and plain-text .txt files, plus pasted text. Older binary .doc files are not supported — open the file in Word or Google Docs and save it as .docx first. For PDFs, use the PDF Citation Generator, which finds the DOI embedded in the file and cites it from the publisher’s official metadata — more accurate than scraping a PDF’s text layer.
- How does the citation cross-check work?
- For author–date styles like APA and Harvard, the tool matches each in-text citation’s first author surname and year against your reference entries. A citation with no matching entry is flagged as missing from the list, and an entry that never appears in the text is flagged as never cited. The check is advisory — always confirm before deleting anything.
- Can I turn the extracted references into formatted citations?
- Yes. Copy the extracted DOIs and paste them into the CitationEasy generator — it accepts a list of DOIs, ISBNs, PMIDs, arXiv IDs, or URLs, one per line, and auto-cites each one in your chosen style. You can also run the extracted list through the APA Citation Checker to catch formatting issues.