Curated, Not Scraped: Why Drug Chatbots Need Experts, Not Just Data
- Mark Tudor, Dr Ivan Ezquerra-Romano
- 3 days ago
- 5 min read

Written by Mark Tudor and Dr Ivan Ezquerra-Romano of Substancy, a UK-based tech organisation specialising in digital solutions and artificial intelligence for the harm reduction sector.
A person opens a chatbot late at night and asks whether the pills they have taken are safe. They may not know what the pills contain. They may be frightened, intoxicated, embarrassed, alone, or avoiding services because of stigma or criminalisation. In that moment, any answer will shape the next decision: seek help, wait, mix substances, or take more. When we talk about drugs, the wrong decision can have fatal consequences.
Why General AI Chatbots Aren't Enough
This explains why AI chatbots that answer questions about drug use cannot be generic chatbots like chatGPT, Gemini, or Claude. These chatbots are trained on vast amounts of data and drug use may or may not be covered in that pile of information. If there's any drug information, it was likely mindlessly scraped from the internet and it is unlikely that drug experts reviewed the content. Data scraping is a collection method, not an evidence standard. It can gather information at scale, but it cannot determine whether that information is current, lawful, relevant, consensual, representative or safe to apply.
By data scraping, we mean the automated collection of online information used to create, train, fine-tune or supply knowledge to a chatbot. The material may come from websites, service pages, forums, social media, news articles, drug-checking alerts, online shops or peer-support communities. The appeal is clear: speed, lower cost and access to emerging information. Drug markets change quickly, while formal guidance can lag behind what people are seeing on the ground.
That need should not be dismissed. Many people who use drugs rely on informal information networks because official systems can be slow, inaccessible, moralising or silent on practical questions. Lived and living experience can be sophisticated harm reduction knowledge, but taking it seriously is not the same as extracting it without consent, context or reciprocity.

The Risks of Scraped Drug Information
The internet contains excellent drug information, but also outdated guidance, duplicated myths, jokes, and marketing claims all dangerous misinformation. A scraped dataset may gather all this together without drug experts’ judgements on how reliable and relevant the data is. A chatbot can then turn that mixture into responses that sound calm, fluent and authoritative.
That false authority is one of the central dangers. Generative AI systems are designed to produce responses, not to guarantee truth. The World Health Organisation (WHO) stresses that AI tools for health must be transparent and subject to appropriate governance. A 2024 study of generative AI responses to real-world substance use and recovery questions found that clinicians often rated answers as high quality. However, some of these answers were dangerous disinformation, including poor responses to suicidal ideation, incorrect emergency helplines and endorsement of home detox. The risk is not just that AI can be wrong. It is that it can be wrong in a voice even experts may trust.
Drug information is unusually context dependent. Risk varies by several factors such as substance, setting, and health. A scraped answer may miss the most important point: the user may not be dealing with the substance they believe they have. The European Union Drugs Agency (EUDA) reported that seven new synthetic opioids were formally notified to the EU Early Warning System in 2024, all of them nitazenes. Information from another country, another year or another drug market may therefore be incomplete or actively misleading.
Scraping can also amplify bias. Online drug information is not evenly produced. Some groups are more visible because they have language access, digital access and relative safety; others are less visible because of poverty, criminalisation, homelessness, surveillance, racism or lack of trust. A chatbot trained mainly on what is easy to collect may reproduce those gaps while appearing to serve everyone.
The Ethical and Governance Challenge
There is also an ethical problem in how community knowledge is used. Online forums and peer networks can contain valuable practical knowledge, but public availability is not the same as consent. A person posting in a drug forum may be seeking support from peers, not agreeing to have their words absorbed into a commercial or institutional AI system.
Legal and governance questions follow. The UK Information Commissioner’s Office has examined web-scraped data in generative AI, including lawful basis, accuracy and individual rights. Copyright and licensing also matter where harm reduction organisations have invested in research and writing. If their work is scraped and used without permission or attribution, trust between those scraping the data and content creators, healthcare providers and service users can be eroded, and the long-term sustainability of producing high-quality harm reduction resources is weakened.
What Responsible Drug AI Should Look Like
None of this means AI chatbots should not be developed for drug information. The better conclusion is that they should not be built as thin conversational layers over ungoverned scraping. Nor should a general-purpose chatbot be treated as a ready-made harm-reduction intervention, with safety delegated to a prompt or disclaimer. A drug-information tool should be a bounded system, designed around defined use cases, approved sources, known exclusions, escalation pathways and ongoing review.
Data for AI systems should be curated by experts rather than scraped. Curated data does not mean ignoring emerging information or excluding lived-experience knowledge. It means information should be selected, checked, dated, attributed and reviewed for a defined purpose before it is allowed to shape advice. Scraping may still have a limited role in horizon scanning, but signals such as new slang or unexpected harms should trigger human review rather than flow directly into automated answers.
Throughout any AI system, human oversight is a must, not a nice-to-have. Researchers, frontliners and people with lived and living experience should all work collaboratively to continuously improve and adapt the system. Drug use and markets evolve and AI systems should too with expert supervision and diverse inputs.
This is not only a challenge for developers. Commissioners, funders, regulators and public-health bodies should not ask only whether a chatbot works, but whether its knowledge base can be inspected, challenged and updated. A system that cannot show where its answers come from, how uncertainty is handled, who reviews its content and what happens when it gets something wrong should not be treated as a credible intervention.
Conclusion
At Substancy, our work has reinforced a practical lesson: the hard part is deciding what a chatbot is allowed to know, when it must stop, and who is responsible for keeping it safe and current.
AI can help widen access to harm reduction information. But in drug contexts, accessibility without governance can become another route to harm. If the knowledge base cannot be examined, challenged and corrected, the system should not be trusted in high risk scenarios. In drug information, the question is not how much the chatbot has read. It is whether anyone can stand behind what it says.




