
The Copyleaks Shared Data Hub is a free, global database of user-submitted documents that Copyleaks compares a submission against, kept separate from the paid Private Cloud Hub that holds one organisation's documents privately. Copyleaks describes it as millions of documents from institutions worldwide, and states that when a document is indexed into it, it becomes available for everyone to compare against.
The question that actually brings people here is narrower: if my school checks my paper against the Shared Data Hub, does my paper end up in it? Reading Copyleaks' own help centre live on 8 September 2026, that question is answered three different ways in three different articles, and the difference between them matters. This guide sets out what the documentation says, where it disagrees with itself, and what a match on your report actually reveals.

What is the Copyleaks Shared Data Hub?
It is one of two document databases Copyleaks offers, and the free one. Copyleaks' API documentation lists them plainly: the Shared Data Hub is a global database containing millions of documents from institutions worldwide, and the Private Cloud Hub is a private database exclusive to your organisation, so that documents stay within your own environment. The documentation labels them in exactly those terms, Shared Data Hub (Free) and Private Cloud Hub (Paid), and an institution can use both at once, storing its documents privately while still comparing them against the shared pool.
The purpose is cross-institutional. Your own school's archive can only catch a student copying from another student at your school. A shared pool catches copying across institutional boundaries, which is the gap it exists to close. That is the same logic behind SafeAssign's Global Reference Database, and the two are worth comparing directly, because they resolve the ownership question in opposite directions.
One naming quirk is worth knowing, because it explains why you will find inconsistent terminology when you search. The feature used to be called the Internal Database. Two of the six help-centre article URLs still contain that old name today, and the body of the article explaining what a match means still refers to "our internal database" in prose. The rename reached the article titles but not the URLs or all of the copy.
Does scanning against the Shared Data Hub also add your paper to it?
This is the consequential question, and Copyleaks' help centre does not answer it consistently. Checked live on 8 September 2026, three articles in the same six-article topic phrase it three different ways.
- "What is the Copyleaks Shared Data Hub?" presents it as a separate decision: "You also have the ability to choose whether or not you would like to submit the file you are scanning into the Shared Data Hub." On that reading, checking against the hub and contributing to it are two independent choices.
- "Does Copyleaks store my content?" presents it as coupled: "Scanning your documents against our Shared Data Hub also means your content will be saved in our secured Database for others to compare scans to in the future."
- "How do I update my settings to save my documents to the Copyleaks Shared Data Hub?" presents it as coupled and automatic: "If you choose to scan our Shared Data Hub, then your content will automatically be saved in the database."
Two of the three say that checking against the hub is what puts you into it. One says the two are separable. That is not a distinction without a difference: it decides whether a student whose institution runs Copyleaks has a paper sitting in a global database indefinitely, or only if someone made an additional choice.
The safe working assumption is the majority reading, that scanning against the hub contributes to it unless the setting says otherwise, because that is what two of three articles state and it is the assumption that fails safely. But the honest answer is that this is a per-institution scan setting your school controls, not something you can determine from the documentation alone, and it is a fair question to put to whoever administers Copyleaks at your institution.

How does a document actually get into the Shared Data Hub?
Through a deliberate indexing step, which is clearer in the API documentation than in the help centre. Copyleaks describes a two-stage process: first you index a document, which stores it without scanning it, using an explicit setting the API calls IndexOnly. Only then do you start the comparison, which runs the indexed documents against each other and against whichever databases were selected.
One constraint on that process is flagged in bold in Copyleaks' own documentation and is easy to miss. Scanning options such as internet matching and AI detection "must be configured during this indexing step", and "cannot be changed later when you start the comparison scan". So the decisions about what a scan will include, including hub participation, are made at upload time rather than revisited when results are produced.
The practical consequence for a student is that this is an institutional configuration decision, made by whoever set up the integration, not something chosen per assignment by an instructor and not something you are prompted about at submission.
What does a Copyleaks Shared Data Hub result mean on a report?
It means your text matched a document stored in that database. Copyleaks states that a result titled "Copyleaks Shared Data Hub" indicates it detected copied content from a document in the database, and that a result reading "Your file [File Name] submitted" means your submission matched something you yourself submitted previously, which is the self-match case that panics people who have legitimately reused their own draft.
What the reader of the report can see is deliberately limited. Copyleaks states that "only identical content will be shown, the rest of the document will be masked with ### to hide any private content". So a match exposes the overlapping passage and nothing else of the other document. It also states that details like the name of the user or organisation that submitted the matched document will not be shared, and separately that no personal data will ever be viewable by others.

What is the difference between an 'Other Files' and a 'Your Files' result?
Copyleaks splits Shared Data Hub results into exactly these two categories, and the privacy difference between them is the single most useful thing on this page for a student.
An "Other Files" result means a user at a different organisation uploaded the document you matched against. The original content is masked to protect that user's privacy, only the identical text is visible, and no user or organisation name is attached. Anonymity here is the documented default.
A "Your Files" result means the match came from inside your own institution, which could be another student or a staff member using your school's learning management system integration. Here the anonymity is conditional. Copyleaks documents an administrator setting called "Identify student information in the results for LMS users", and states that when it is enabled, instructors and admins can see identifying information about the student who was copied from. If it is not enabled, the report shows only the label "Your Organization's File".
So the honest answer to "can someone see who I am from a match" is: not across institutions, and inside your own institution it depends on a setting your administrator controls. That is a meaningfully different answer from a flat yes or no, and it is the sort of thing worth knowing before assuming a report is anonymous.
Can you delete a document from the Copyleaks Shared Data Hub?
Yes, and this is where Copyleaks differs sharply from the closest comparable system. Deletion is clearly possible: one help-centre article states that "files can be permanently deleted from the Shared Data Hub" and links out to do it, and another states that "you always have the option to delete the documents you have submitted from the Shared Data Hub".
The route is less settled than the principle. A third article says that to permanently remove previously submitted content you should contact support. So across three articles in the same topic there is a self-service link, a general assurance, and a support request, without a clear statement of which applies to whom. If you need something removed, expect that the answer in practice may depend on whether you submitted it yourself or your institution submitted it on your behalf through an integration.
It is also worth being realistic about who holds this control. If your work entered the hub through your university's Copyleaks integration, the account that submitted it is your institution's, not yours, so the request will usually go through your school rather than directly to the vendor.
How does this compare with SafeAssign's Global Reference Database?
They solve the same problem and take opposite positions on getting back out, which is the comparison worth having if your school uses one and you have read about the other.
- Contribution. Blackboard frames the Global Reference Database as a voluntary donation, where students choose to submit copies of their papers, and it is deliberately kept separate from an institution's own internal archive. Copyleaks frames hub participation as a scan setting, and two of its three help articles describe contribution as following automatically from checking against the hub.
- Deletion. This is the sharp difference. Blackboard states that students who contribute agree not to delete those papers later. Copyleaks states the opposite, that files can be permanently deleted from the Shared Data Hub.
- Identification. Copyleaks documents an explicit admin setting that can attach a student's identity to a within-institution match. That level of detail is unusual to find stated openly, and it is genuinely useful.
The SafeAssign side of this comparison, including the size of that archive and what agreeing not to delete actually commits you to, is set out in what the Global Reference Database is.
And if the term on your report is the broader document rather than the database behind it, what an originality report actually contains covers how to read one.
Is the Shared Data Hub an AI detector?
No, and conflating the two produces most of the confusion around Copyleaks reports. The Shared Data Hub is text-to-text matching: your words are compared against stored documents, and a match arrives with a source attached that a person can inspect. Copyleaks' AI detection is a separate statistical judgement about writing style, with no source to check, and a document runs through that scan independently of whether hub matching is switched on.
The two produce unrelated results on the same submission. Original writing that happens to reuse a common phrase can raise a small match while scoring zero for AI. Text generated this morning will usually match nothing at all, because it is genuinely novel, while scoring high on the AI indicator. A clean similarity score is not evidence about AI, and a clean AI score is not evidence about copying.
How reliable Copyleaks' AI side is, including what happens to its accuracy once text has been edited, is covered separately in does Copyleaks detect humanized AI text.
The underlying mechanism, and why a statistical detector cannot show you a source the way a matching database can, is explained in how ChatGPT detectors work.
What should you check before you submit?
Four things, and all of them are answerable before a report exists rather than after one has gone wrong.
- Whether your institution's Copyleaks configuration contributes submissions to the Shared Data Hub. This is a scan setting your administrator controls, it is decided at upload time, and it is a reasonable question to ask.
- Whether the "Identify student information in the results" setting is enabled, if you care whether a within-institution match can name you.
- Whether reusing your own earlier work is permitted on this assignment. A "Your file submitted" self-match is a documented result type, and it catches people who legitimately built on a previous draft. Self-plagiarism rules vary by institution and by course.
- What a match actually means before you react to one. It is a passage with a source attached, not a verdict, and the masked context means you are seeing the overlap rather than the whole picture.
If what you actually want to know is how your writing scores on the AI side before anyone else sees it, our free AI detector gives you a reading in about thirty seconds with nothing to install.
And where your own original writing reads as uniform enough to trip that indicator, our AI humanizer rebuilds sentence rhythm rather than swapping vocabulary, which is the property those detectors actually score.
How this was made: every quotation and figure above was read directly in the browser on 2026-09-08 from Copyleaks' own API documentation and help centre, not from search snippets or third-party summaries, and not via an automated page fetch, which returned nothing usable on the help-centre pages because they are JavaScript-rendered. The free-versus-paid labelling, the two-step IndexOnly process and the note that scan options cannot be changed after indexing come from the Data Hubs page of docs.copyleaks.com. The three conflicting descriptions of whether scanning contributes to the hub, the three deletion routes, the ### masking, the "Other Files" and "Your Files" categories and the "Identify student information in the results" setting come from four separate articles in the help centre's Shared Data Hub topic. Ours rather than Copyleaks' are the observations that the topic contains exactly six articles, that its "what is it" article carried 7,467 views against 8,383 for the other five combined when counted on 2026-09-08, that the three articles disagree, and that two article URLs plus one article body still carry the feature's former name. View counts change, which is why that reading is dated rather than presented as a fixed number. We have not run a scan inside a paid institutional Copyleaks account and claim no finding about one, and no accuracy percentage for Copyleaks appears in this post because none was re-verified this cycle. Drafting is AI-assisted, with a person verifying each claim against its live source before publishing.
Frequently asked questions
What is the Copyleaks Shared Data Hub?
It is Copyleaks' free, global database of user-submitted documents, described by Copyleaks as millions of documents from institutions worldwide, which submissions can be compared against to find matches across institutional boundaries. It sits alongside a paid Private Cloud Hub that keeps one organisation's documents private, and an institution can use both at once.
Does Copyleaks store my content in the Shared Data Hub?
It can, and Copyleaks' own help centre is inconsistent about when. One article says you choose separately whether to submit a file you are scanning; two others say that scanning against the hub means your content is saved automatically. Checked on 8 September 2026, the majority reading is that checking against the hub contributes to it, but it is governed by a scan setting your institution controls.
Can I delete my document from the Copyleaks Shared Data Hub?
Yes. Copyleaks states that files can be permanently deleted from the Shared Data Hub, and separately that you always have the option to delete documents you have submitted. A third article directs you to contact support for permanent removal. If your work was submitted through your university's integration, the request usually has to go through your institution, since the submitting account is theirs.
Can someone see my name from a Shared Data Hub match?
Not across institutions. Copyleaks states that details like the name of the user or organisation that submitted a matched document will not be shared, and that only identical content is shown while the rest is masked. Inside your own institution it depends on an administrator setting called "Identify student information in the results for LMS users", which when enabled lets instructors and admins see who was copied from.
What does 'Your file submitted' mean on a Copyleaks report?
It means your submission matched a document you submitted yourself previously, rather than someone else's work. Copyleaks documents this as a distinct result type. It commonly appears when a student reuses their own earlier draft, and whether that is acceptable is an institutional self-plagiarism question rather than a technical one.
Does the Shared Data Hub detect AI or ChatGPT?
No. The Shared Data Hub is a plagiarism-matching database that compares your text against stored documents and returns matched passages with sources attached. Copyleaks' AI detection is a separate statistical scan of writing style that runs independently of any hub setting, and the two produce unrelated results on the same submission.
Is a Shared Data Hub match proof of plagiarism?
No. A match identifies a passage that is identical to text in a stored document and attaches a source to it. Common phrases, correctly quoted material, shared assignment prompts and your own earlier work all produce matches. It is evidence a person has to interpret, not a verdict, and the masked context means the reader is seeing the overlap rather than the full document.
Does every school using Copyleaks have the Shared Data Hub enabled?
No. Hub participation is an institutional scan setting rather than an automatic part of every Copyleaks scan, and Copyleaks' documentation shows it is configured at the point a document is indexed. Whether your school compares against the shared hub, contributes to it, or uses a private hub instead varies by institution.


