Module 6 — Autonomous Scientific Intelligence System
You are not building a summarizer. You are building a six-agent AI research department
that understands papers, detects contradictions, generates hypotheses, maps research gaps,
identifies commercial opportunities, and writes publication-ready reports — all from a single user
query in Google AI Studio.
What You Build
SciIntel Engine
6-agent autonomous research system: Literature
Analyst, Evidence Validator, Hypothesis Generator, Gap Detector, Innovation Strategist,
Scientific Writer.
What You Learn
5 Core Techniques
Scientific Specification, Agent Persona
Engineering, Evidence Chain Prompting, Hypothesis Prompting, and Research Gap Mapping.
Platform
Google AI Studio
Free. No credit card. Sign in with Google. Use
Build mode for the app + Gemini 2.0 Flash for the agents. Upload PDFs directly.
Your Output
Live Shareable App
A shareable research intelligence URL. Upload
any papers, get structured intelligence output across 5 levels of scientific reasoning.
Why This Is Different From a Summarizer
Every researcher has access to Google Scholar and ChatGPT. They can already summarize papers.
What they cannot do easily is:
1. Detect contradictions across 10 papers simultaneously and explain which study design caused the disagreement
2. Generate novel hypotheses by identifying unexplored intersections between two research streams
3. Map a research gap heatmap — what has been heavily studied vs what is almost entirely unexplored
4. Produce a grant proposal first draft with aims, rationale, and novelty statement from the literature
5. Identify therapeutic opportunity spaces that no single paper has articulated
This system does all five. That is why it is an intelligence engine — not a summarizer.
1. Detect contradictions across 10 papers simultaneously and explain which study design caused the disagreement
2. Generate novel hypotheses by identifying unexplored intersections between two research streams
3. Map a research gap heatmap — what has been heavily studied vs what is almost entirely unexplored
4. Produce a grant proposal first draft with aims, rationale, and novelty statement from the literature
5. Identify therapeutic opportunity spaces that no single paper has articulated
This system does all five. That is why it is an intelligence engine — not a summarizer.
The App You Are Building — SciIntel Engine
Upload & Analyze
Evidence Map
Hypotheses
Gap Heatmap
Strategy
Reports
6
Papers Loaded
3
Contradictions
12
Hypotheses
AGENT 1
— Literature Analyst · Active
Extracted 6 objectives, 4
distinct methodologies (RCT n=312, observational n=1,847, meta-analysis 18 studies,
in-vitro). Key finding: semaglutide shows histological improvement in NASH
(p<0.01, SUSTAIN-6 subgroup).
AGENT 2
— Evidence Validator · Contradiction Detected
Study A (2021) reports 42%
reduction in liver fat. Study D (2023) reports 18% reduction. Same drug class,
different endpoint definitions. Study A used MRI-PDFF; Study D used liver biopsy NAS
score. Methodological disagreement — not biological.
Certification Path: Complete all 6 build phases, score 6/10 on the
quiz, and build all 6 agents to unlock your certificate. Each phase teaches a specific prompting
technique through the act of building.
The Vision — 5 Levels of Scientific Intelligence
Most AI tools stop at Level 1. Your system goes to Level 5. Each level is a distinct
capability — built through a different prompting technique — that stacks on top of the previous one.
Level 1 — Understands
Literature Analyst Agent
Reads uploaded papers and
extracts: objectives, methodology, sample sizes, primary endpoints, key findings,
statistical significance, and study limitations. Output is structured — not paragraphs.
Every finding is traceable to a specific paper and page.
Technique used: Scientific
Specification Prompting — defining exact extraction fields with output format constraints
before Gemini reads a single word.
Level 2 — Analyzes
Evidence Validator Agent
Compares conclusions across all
uploaded papers simultaneously. Flags contradictions, identifies the source of disagreement
(different populations, different endpoint definitions, different statistical methods),
grades evidence strength (Level A/B/C), and maps scientific consensus vs controversy.
Technique used: Evidence Chain
Prompting — forcing the agent to trace every claim back to its source study and grading it
before any synthesis is allowed.
Level 3 — Thinks
Hypothesis Generator + Gap
Detector Agents
Goes beyond what is in the
papers to what is missing. Identifies unexplored intersections between research streams,
generates novel testable hypotheses with mechanistic rationale, and builds a research gap
heatmap — showing which areas have abundant evidence, which are emerging, and which are
almost entirely unexplored.
Technique used: Hypothesis Prompting
— constraining the agent to generate only falsifiable, mechanistically grounded hypotheses
with experimental designs to test them.
Level 4 — Strategizes
Innovation Strategist Agent
Translates scientific findings
into strategic intelligence: which findings have commercial potential, which disease areas
are underserved, what combination therapies are suggested by the literature but untested,
which findings could support a grant proposal, and what publication strategy maximizes
impact.
Technique used: Research Gap Mapping
— a structured two-axis analysis (evidence strength x commercial potential) that generates
strategic recommendations, not just summaries.
Level 5 — Acts Like a Scientific Think Tank
Executive Scientific Writer
Agent
Synthesizes all 5 agent outputs
into publication-ready documents: structured literature reviews in journal format, journal
club presentation summaries, grant proposal drafts (specific aims, rationale, novelty),
research roadmaps, and innovation opportunity maps. Every output is formatted for immediate
use — not a starting draft that needs major editing.
Technique used: Agent Persona
Engineering — building a writer persona with journal-specific style rules, citation format
constraints, and self-critique loops that catch compliance failures before delivery.
The Single Query That Drives All 5 Levels
A researcher uploads 6 papers
on GLP-1 receptor agonists in NASH and types one query. Here is what the system produces —
automatically, sequentially, across all 6 agents:
USER QUERY: "Analyze the current evidence for GLP-1 receptor agonists in NASH.
What is the strongest finding, where do studies disagree, what is missing, and should my lab
pursue this direction?"
AGENT 1 OUTPUT: Structured extraction table — 6 papers, 12 findings, 4 methodologies graded
AGENT 2 OUTPUT: 3 contradictions detected — root cause: endpoint definition differences
AGENT 3 OUTPUT: 5 novel hypotheses — dual GLP-1/FGF21 mechanism in resistant cases
AGENT 4 OUTPUT: Research gap heatmap — pediatric NASH, combination therapy, biomarker prediction
ALL red (unexplored)
AGENT 5 OUTPUT: Strategic recommendation — pursue pediatric NASH direction, 2 grant
opportunities identified
AGENT 6 OUTPUT: 3-page literature review + specific aims section for R01 grant application
What just happened: One 15-word query triggered a complete
research intelligence workflow. The researcher now has the equivalent of a full literature
review team's output in structured, actionable format. That is the power of the system you are
building.
Your 6 Specialized Research Agents
Each agent has a specific scientific identity, a defined scope, and behavioral rules
that prevent it from overstepping its mandate. Click any agent to see its full system prompt
template and the prompting technique that powers it.
Why 6 Agents Instead of 1?
Single-Agent Approach
"Analyze these papers and tell me everything important."
— Gemini gives a 600-word summary. Generic. No structure. Mixes analysis with opinion.
Cannot cite specific contradictions. No hypotheses. No strategy.
Result: A slightly better Wikipedia summary.
Multi-Agent Approach
Agent 1 extracts ONLY findings — nothing else.
Agent 2 compares ONLY contradictions — nothing else.
Agent 3 generates ONLY hypotheses — with constraints.
Agent 4 maps ONLY gaps — with evidence grading.
Agent 5 identifies ONLY strategy — with commercial lens.
Agent 6 writes ONLY reports — in journal format.
Result: Six specialized, high-precision outputs that stack into a complete research
intelligence brief.
The core principle: Specialization produces higher quality
than generalization. Each agent knows exactly what it must do, what it must never do, and what
format its output must take. This is how real research teams work — not one generalist, but six
specialists.
Build Your SciIntel Engine — 6 Phases
Each phase builds one part of your system and teaches one prompting technique through
the act of building it. Follow the phases in order — each one depends on the previous.
10 Power Features of Your System
These are the advanced capabilities that make your system feel like a scientific think
tank — not a chatbot. Each feature is built through a specific prompting pattern explained below.
Feature 01 — Contradiction Detector
Finds conflicting conclusions across papers
The system compares
endpoint definitions, study populations, and statistical methods across all uploaded papers
— then identifies where two studies report different conclusions on the same question and
explains the methodological reason for the disagreement.
[ROLE] Evidence Validator — you find contradictions, not consensus.
[TASK] Compare all uploaded papers on [TOPIC]. Identify every instance where two or more
studies report conflicting conclusions on the same question.
[FOR EACH CONTRADICTION, OUTPUT]
Contradiction [N]:
Claim A: "[exact finding]" — Source: [Paper title, year, page N]
Claim B: "[exact finding]" — Source: [Paper title, year, page N]
Root cause: [Endpoint definition difference / Population difference / Statistical method
/ Timepoint difference]
Implication: [What this contradiction means for clinical practice]
Resolution: [Which study is more reliable and why — or "unresolved, needs head-to-head
trial"]
[CONSTRAINTS]
NEVER synthesize contradictions into consensus — your job is to find the disagreements
NEVER report a contradiction unless both claims are explicitly stated in the papers
Grade each contradiction: MAJOR (different clinical conclusions) or MINOR (different
magnitude, same direction)
Feature 02 — Novel Hypothesis Engine
Generates falsifiable scientific hypotheses
Forces the agent to
generate only hypotheses that are: grounded in the literature, mechanistically explained,
falsifiable with a specific experimental design, and novel — not already tested in the
uploaded papers.
[ROLE] Hypothesis Generator — you generate scientific ideas, not
summaries.
[TASK] Based on the gaps and patterns in the uploaded literature, generate [N] novel,
testable hypotheses.
[EACH HYPOTHESIS MUST CONTAIN]
H[N]: [One sentence hypothesis in "If X, then Y, because Z" format]
Mechanistic rationale: [Why this is biologically/chemically/clinically plausible — 2-3
sentences]
Evidence basis: [Which specific findings from the literature support this direction]
What is missing: [What the current literature fails to test that this hypothesis
addresses]
Experimental design to test it: [Specific study type, endpoints, sample, duration]
Novelty score: [HIGH = not tested anywhere / MEDIUM = tested in different context / LOW
= extension of existing work]
[HARD CONSTRAINTS]
NEVER generate a hypothesis already tested in the uploaded papers
NEVER generate a hypothesis without a mechanistic rationale
NEVER use vague language — "may" is not acceptable; use "will" or "is predicted to"
Flag any hypothesis where the evidence basis is weak: [WEAK EVIDENCE — speculative]
Feature 03 — Research Gap Heatmap
Visual map of explored vs unexplored territory
The agent maps all
sub-topics of the research field on two dimensions — evidence density (how much has been
studied) and clinical importance (how much it matters). Outputs a structured table that can
be rendered as a color-coded heatmap.
[ROLE] Research Gap Detector — you find what is missing, not what exists.
[TASK] Map the research landscape for [FIELD/TOPIC] based on the uploaded literature.
[OUTPUT FORMAT — HEATMAP TABLE]
| Sub-topic | Evidence Density | Clinical Importance | Gap Score | Priority |
|-----------|-----------------|---------------------|-----------|----------|
Each sub-topic: 1-3 words. Evidence density: HIGH / MEDIUM / LOW / NONE. Clinical
importance: HIGH / MEDIUM / LOW. Gap score: (Clinical importance) x (inverse of evidence
density) — scale 1-10. Priority: URGENT / HIGH / MEDIUM / DEPRIORITIZE.
[AFTER THE TABLE]
Top 3 Research Priorities: [Sub-topics with highest gap scores]
Quick wins (LOW evidence + FAST to study): [Sub-topics where a single well-designed
study could fill the gap]
Long-term opportunities (NONE evidence + HIGH importance): [Emerging areas worth a
research program]
[SCOPE]
Include at minimum: pediatric populations, combination therapies, biomarker development,
underserved geographies, long-term outcomes (5+ years), specific patient subgroups not
studied
Feature 04 — Grant Proposal Assistant
First-draft grant sections from literature
Generates publication-ready
Specific Aims, Significance, and Innovation sections for an NIH/ICMR grant proposal —
grounded in the uploaded literature, using the exact language and structure that review
panels look for.
[ROLE] Executive Scientific Writer — Grant Proposal Specialist. You write
in NIH R01 format. Every claim must be supported by a citation from the uploaded
literature.
[GRANT TARGET] [NIH R01 / ICMR / Wellcome Trust / DBT — specify]
[PI INSTITUTION] [Your institution]
[RESEARCH FOCUS] [1-sentence description]
[TASK] Write the following grant sections based on the uploaded literature:
SECTION 1 — SPECIFIC AIMS (1 page max)
Structure: Opening paragraph (the problem) + Aim 1 + Aim 2 + Aim 3 + Closing paragraph
(impact)
Each Aim: "[Verb] [specific outcome] using [methodology] in [population] to [test
specific hypothesis]"
Do NOT start aims with "To investigate" or "To explore" — these are too vague for R01
review.
SECTION 2 — SIGNIFICANCE (0.5 page)
Paragraph 1: Current state of evidence — cite 3 strongest findings from uploaded papers
Paragraph 2: Critical gap — what is not known, grounded in the gap analysis
Paragraph 3: Impact statement — how this research changes clinical practice
SECTION 3 — INNOVATION (0.5 page)
What is novel about this approach (3 bullet points, each with a citation showing what
currently exists)
Why existing approaches are insufficient (2 sentences, specific not generic)
[OUTPUT CONSTRAINTS]
Present tense only for established facts, future tense for proposed work
NEVER use: "cutting-edge", "novel approach", "important study" — show innovation, don't
claim it
Every factual statement must have [Author, Year] in-text citation
Feature 05 — Journal Club Generator
Automatic journal club presentation from a
single paper
Generates everything needed
to present a paper at a journal club: structured summary, critical evaluation, faculty
questions, audience questions, clinical relevance, and a verdict on whether the findings
should change practice.
[ROLE] Journal Club Moderator with 20 years of academic medicine
experience. You present papers critically — you never describe without evaluating.
[PAPER] [Upload the paper or paste abstract + methods]
[JOURNAL CLUB OUTPUT — EXACT STRUCTURE]
1. PAPER IN 3 SENTENCES
What they did / What they found / Why it matters (or doesn't)
2. STUDY DESIGN SCORECARD
Population representativeness: [1-5] — [reason]
Endpoint validity: [1-5] — [reason]
Statistical rigor: [1-5] — [reason]
Conflict of interest risk: [1-5] — [reason]
Overall grade: [A/B/C/D]
3. WHAT THEY DID WELL (2 points max)
4. WHAT THEY DID POORLY (2 points max — be specific, not vague)
5. THE QUESTION THEY DIDN'T ASK (1 specific gap this paper ignores)
6. FACULTY QUESTIONS (3 questions a senior faculty member would ask — not softballs)
7. SHOULD THIS CHANGE PRACTICE?
Verdict: [YES / NO / ONLY IN SPECIFIC SUBGROUP — specify]
Reason: [2 sentences, specific to this paper's evidence quality]
Feature 06 — Citation Intelligence
Identifies the most and least influential
papers
Ranks uploaded papers by
their influence on the field — identifying foundational studies every researcher must cite,
emerging studies that will define the next generation of work, and weakly supported papers
that should not be used as sole evidence.
[ROLE] Citation Intelligence Analyst. You evaluate the relative
importance of studies, not their content.
[TASK] From the uploaded papers, classify each by its role in the field:
TIER 1 — FOUNDATIONAL (must cite in any paper on this topic)
Paper: [Title, Author, Year]
Why foundational: [What this paper established that changed the field]
Citation status: [Landmark / Widely replicated / Standard reference]
TIER 2 — EMERGING (defines the next generation of research)
Paper: [Title, Author, Year]
Why emerging: [What new direction this paper opens]
Likely trajectory: [Will become Tier 1 if replicated / Needs validation first]
TIER 3 — SUPPORTING (valid but limited scope)
Paper: [Title, Author, Year]
Appropriate use: [When to cite this / When NOT to cite this]
TIER 4 — WEAKLY SUPPORTED (use with caution)
Paper: [Title, Author, Year]
Weakness: [Small n / Single center / Non-peer-reviewed / Unfunded replication lacking]
Warning: [Never cite as sole evidence for X]
[FINAL NOTE]
Identify 1 paper from the uploaded set that is most likely to be cited in the next major
systematic review on this topic. Explain why.
Using these features: Each prompt above is a
ready-to-use agent prompt. Copy it, paste it as the System Instructions in Google AI Studio, then
upload your papers and ask your research question. The agent will execute the exact analysis defined
by the prompt.
Debug & Upgrade Your Research System
When your agents produce disappointing output, it is almost always a prompt design
problem — not a Gemini limitation. Here are the 6 most common failures and the exact fixes.
Problem 01 — Agent Gives Summaries Instead of Analysis
Gemini describes what the
papers say instead of analyzing, comparing, or critiquing them.
You are summarizing instead of analyzing. Stop.
Your role is NOT to describe what each paper says.
Your role IS to compare papers against each other and identify patterns, contradictions,
and gaps.
Restart with this constraint:
BEFORE writing any output, answer: "What do at least 2 of these papers DISAGREE on?"
If you cannot find a disagreement, you are not looking hard enough — every research
field has one.
Your output must contain at minimum:
1. One comparison across papers (not a description of one paper)
2. One contradiction or tension you identified
3. One thing NONE of the papers address
Do not restate what is already in the papers. Add analytical value beyond what is there.
Problem 02 — Hypotheses Are Too Obvious
The Hypothesis Generator
produces ideas that are already tested in the uploaded papers or are so obvious they add no
value.
The hypotheses you generated are already tested in the literature I
uploaded. That is not what I asked for.
A novel hypothesis is one where NONE of the uploaded papers have tested it.
Apply these three filters to every hypothesis you generate:
Filter 1 — NOVELTY CHECK: Search your output for the words "has been shown", "studies
demonstrate", "evidence suggests". If they appear, that hypothesis describes existing
evidence — delete it.
Filter 2 — INTERSECTION TEST: The most powerful novel hypotheses come from combining two
research streams that have never been connected. Identify two topics in the uploaded
papers that share a biological mechanism but have never been studied together.
Filter 3 — FALSIFIABILITY TEST: If you cannot describe a specific experiment that would
prove the hypothesis wrong, it is not a scientific hypothesis — it is an opinion. Delete
it.
Regenerate [N] hypotheses that pass all three filters.
Problem 03 — Evidence Grading Is Too Generous
The Evidence Validator
gives Level A evidence to observational studies or small pilot trials, inflating confidence
in weak findings.
Your evidence grading is too generous. Apply the Oxford Centre for
Evidence-Based Medicine (OCEBM) hierarchy strictly:
LEVEL 1A: Systematic review of RCTs with homogeneity
LEVEL 1B: Individual RCT with narrow confidence interval
LEVEL 2A: Systematic review of cohort studies
LEVEL 2B: Individual cohort study or low-quality RCT
LEVEL 3A: Systematic review of case-control studies
LEVEL 3B: Individual case-control study
LEVEL 4: Case series, poor quality cohort or case-control
LEVEL 5: Expert opinion, animal studies, in-vitro only
Re-grade every finding in your previous output using this hierarchy.
For each finding, state: [OCEBM Level N] [Study type] [n=sample size] [confidence
interval if reported]
Downgrade any finding that lacks: randomization (for RCTs), pre-registration, or sample
size justification.
Flag any finding where the conclusion exceeds what the evidence level can support:
[OVERCLAIM DETECTED]
Problem 04 — Literature Review Sounds Generic
The Executive Writer
produces a review that could have been written without reading the papers — full of generic
sentences any AI would produce.
This literature review is too generic. A reviewer could not tell which
papers I uploaded from reading your output.
Apply the Specificity Test: every sentence in the review must be traceable to a specific
paper, finding, or data point from my uploaded documents.
Rewrite with these rules:
RULE 1: Every paragraph opens with a specific finding — not a general statement.
"Semaglutide at 2.4mg weekly reduced liver fat by 34% (p=0.003, n=312, LEAN trial 2021)"
NOT "GLP-1 agonists have shown promising results."
RULE 2: Every transition word ("however", "conversely", "importantly") must be followed
by a specific contradiction or nuance from the papers — not a generic modifier.
RULE 3: The conclusion paragraph must contain the one thing this body of literature has
failed to study — specific, not "more research is needed."
RULE 4: Remove every phrase that could appear in a review of any other topic:
"significant advances have been made", "growing body of evidence", "further studies are
warranted."
Rewrite the review with these 4 rules applied.
Problem 05 — Strategy Agent Is Not Strategic
The Innovation Strategist
gives generic "this could be commercially valuable" statements without specific direction or
actionable recommendations.
Your strategic output is generic. "This could be commercially valuable"
is not strategy.
Apply the OGSM Framework to every strategic recommendation:
O — Objective: What specific outcome does this strategy achieve? (1 sentence,
measurable)
G — Goal: What is the specific metric that defines success? (a number or milestone)
S — Strategy: What specific approach achieves the goal? (not "research more" — specific
actions)
M — Measure: How will we know in 12 months if this worked? (specific leading indicator)
For each opportunity you identified, apply OGSM:
Opportunity: [Name]
O: [Specific objective]
G: [Measurable goal — e.g. "First-in-class IND filing within 24 months"]
S: [Specific strategy — e.g. "License the pediatric indication from [Company] while
building own Phase 2 data in the parallel biomarker cohort"]
M: [12-month leading indicator]
Delete any recommendation where you cannot fill in all 4 fields with specifics.
Problem 06 — Gemini Hallucinated Citations
The agent cited papers that
were not uploaded, invented author names, or stated findings that do not appear in any of
the uploaded documents.
WARNING: You have cited sources that were not in my uploaded documents.
This is a hallucination — it is the most serious error in scientific AI.
Apply Citation Discipline Mode immediately:
RULE 1: You may ONLY cite papers that are in my uploaded documents. If you cannot find
the paper in my uploads, do not cite it.
RULE 2: Every citation must include the exact sentence or data point from that paper
that supports your claim. If you cannot quote it, you do not cite it.
RULE 3: If you want to make a point that is not supported by my uploaded papers, write:
[EXTERNAL KNOWLEDGE — not from uploaded documents] before the statement.
RULE 4: If you are uncertain whether a finding is from my documents or your training
data, write: [CITATION UNCERTAIN — verify manually].
Rewrite your output with these 4 rules. Delete any claim that cannot meet Rule 2. This
is non-negotiable — a hallucinated citation in a scientific document is worse than no
citation.
Final Quiz — 10 Questions
Test your mastery of scientific intelligence prompting. Score 6/10 or above to unlock
your certificate. Each question tests a concept you built with — not just memorized.
Score:
0
/ 10
Your Certificate
Complete all 6 phases, score 6/10 on the quiz, and build all 6 agents to unlock your
certificate.
Phases Done
0/6
Quiz Score
0/10
Agents Built
0/6
Certificate
Locked