SYSTEM NOTICE

Auto translation by AI. Be sure, accuracy, nuances and authorial intent may not be fully reflected.
見出し画像

A Savior for Paper Screening! Dramatically Streamline Your Literature Review with Elicit

Elicit: An AI Tool to Accelerate Systematic Reviews

For researchers, creating systematic reviews, literature reviews, and clinical practice guidelines is a time-consuming and labor-intensive task. In this context, Elicit is gaining attention as an innovative tool that leverages AI to dramatically streamline these processes.

In particular, its ability to significantly reduce the time required for literature search and screening is appealing, leading to a dramatic improvement in research productivity. I recently had a project that required reviewing a large number of papers and found it extremely useful, so I will delve into the appeal of Elicit along with my personal impressions of using it, and also explore comparisons with OpenAI's Deep Research and the potential for integration with LLMs (Large Language Models).

Summary of this article

What Elicit can do

Elicit supports the systematic review process through the following steps:

1. Paper Collection

  • Entering Research Queries: Enter questions related to your research topic (e.g., "How does cognitive behavioral therapy affect depression?"). The AI will refine the query as needed.

  • Adding Papers: Upload PDFs you have on hand, select from your Elicit library or Zotero, or search for and add relevant papers from Elicit's database of over 125 million papers based on Semantic Scholar. It supports up to 1,000 papers.

  • Points to Note: Whether it directly covers specific databases such as PubMed, MEDLINE, and the Cochrane Library is not explicitly stated in official information. Therefore, for research requiring comprehensiveness, such as systematic reviews, it is recommended to consider using it in conjunction with these databases.

2. Screening (Narrowing down literature)

  • Automatic Generation/Customization of Screening Items: Elicit suggests items such as "Is this study a randomized controlled trial?". Users can modify these or add their own unique items.

  • Pilot Screening: Test with a sample of about 100 papers to adjust the criteria.

  • Automatic Screening and Scoring: The AI scores the degree of match (e.g., 4 points or higher is almost certainly accepted, 3-point range requires caution, and there are papers worth picking up even in the 2-point range).

  • Setting Cutoff Values and Manual Checks: You can change the score criteria yourself to narrow down the accepted papers, while also manually adding or excluding papers above or below the cutoff score.

Since it automatically generates screening items, you can modify them or create new items yourself.
For each screening item, it shows whether the paper matches.
By changing the score threshold, you can adjust the number of papers to proceed to data extraction.
It is better to manually check and override the acceptance/rejection of papers around the score threshold individually.

3. Data Extraction (Up to 200 items)

Extract data such as study design and results from screened papers.You can also set the items to be extracted yourself. However, the Pro plan has a limit of 200 items at a time and 200 items per month. The Team plan ($79/month) can be expanded to 300 items per month.Once you start extraction, this quota will be consumed, so please decide on your items carefully before starting. There are also errors in data extraction. For example, results from multivariate analysis and univariate analysis might be listed together, so verification is necessary. However, if you click the * mark after the data, you can jump to the corresponding section in the paper, allowing for quick verification.

Clicking the * mark allows you to jump to the corresponding section in the paper
Pricing Plans

4. Report Creation (based on up to 40 papers)

  • Automatically generate detailed reports with evidence based on the extracted data.

  • Note that even if you target 100 papers for data extraction, the report will only cover 40. It feels a bit like a loss, and because of this 40-paper limit, you cannot leave everything to it to write a systematic review.

  • Reports can be exported as PDF or links can be shared with others.

  • There are still some errors in the reports, so verification is necessary. It sometimes cites a description of paper B, which is referenced within paper A, as a claim made by paper A (though this is a fairly common mistake even when done by humans).


What's Great About Elicit!

  • Overwhelming time savings: What used to take days for literature screening during guideline creation or reviews can be done in a few hours using Elicit. It is overwhelmingly faster than reading abstracts one by one yourself.

  • Flexible and convenient screening: It is convenient to be able to customize it yourself based on the screening items proposed by the AI, and adjust while checking the results in a pilot. A score of 4 or higher is almost certainly acceptable, but scores in the 3s can include questionable papers, and since there are usable ones even in the 2s, manual checking is necessary to prevent omissions.

  • High accuracy but not perfect: AI screening and data extraction are quite accurate, with an experienced accuracy of about 95%. However, papers with similar target diseases or different interventions despite having the same target may be mixed in, so it is best to always check near the cutoff value to avoid missing important papers.

  • LLM Integration via CSV Export: It is fantastic that you can output extracted data as a CSV. Since you can feed this into an LLM to assist with writing, you can move seamlessly from review to paper drafting.

  • A Savior for Clinical Guideline Development: I am impressed by how much easier screening and evidence organization have become for guideline development, which involves handling a vast amount of literature.

  • Looking Forward to Evolution: I used the data extraction feature a year ago, and the accuracy has improved significantly. Elicit is evolving, so I am excited about its future.


Comparison with OpenAI's Deep Research

While OpenAI's Deep Research provides excellent results through advanced search and analysis using o3, challenges remain regarding the comprehensiveness of literature. On the other hand, Elicit excels in covering a wide range of literature thanks to its database of over 125 million items and flexible screening features. It is particularly strong in situations where you have to handle a large number of papers.

LLM Integration

Furthermore, Elicit's CSV export feature is a major strength. If you compile the results of screening or data extraction into a CSV and input them into an LLM (ChatGPT, Gemini, etc.), you can seamlessly proceed to summarizing papers, drafting, or even writing review articles. Finishing a review article with an LLM using data obtained from Elicit's strength in comprehensiveness—that workflow is becoming a reality. (In practice, due to context window limitations, it is better to reduce the number of input tokens by trimming items in the CSV.)


Limitations of Elicit

  • Database Comprehensiveness: Since Elicit is based on the Semantic Scholar database, it does not cover all papers included in specific databases like PubMed, MEDLINE, or the Cochrane Library. When comprehensiveness is critical, you may need to use these databases in parallel and perform manual searching and screening.

  • Initial Screening Limit: The number of papers initially displayed in Elicit seems to be capped at around 500. Therefore, when dealing with large-scale search results, manual processing or integration with other tools may be necessary.

  • Data Extraction Limits: 200 items at a time, and 200 items per month (once you hit 200, you are done for the month). The Team plan increases this to 300 items per month, but this may still be insufficient in fields with many relevant papers.

  • Report Generation: Supports up to 40 items. Other methods are required for large-scale reviews.

  • Accuracy: The precision of extracted data is approximately 95% (anecdotal). Manual verification is essential.Open-access papers have their full text in Elicit's database, but some only have abstracts, which may also contribute to differences in accuracy.


Summary

Elicit is a reliable AI tool that streamlines systematic reviews, literature reviews, and the creation of clinical practice guidelines. By combining AI automation with manual checks, you can produce high-quality results in a short amount of time. The ability to complement the comprehensiveness of OpenAI/Deep Research and connect to writing through integration with LLMs greatly expands the possibilities of research. Although there are limitations, for researchers looking to significantly reduce time and effort, Elicit is truly a savior.


いいなと思ったら応援しよう!