OSINT — Hands-on lab
You take a real target — a public bug bounty program — and you take it apart in fully passive mode. Deliverable: a complete target sheet, with five ranked priority attack points.
Plan for: 2 to 3 hours.
Deliverable: a complete ~/osint/<target>/ folder, with a 2-page fiche-cible.md sheet at the top.
You must pick a public bug bounty program whose scope explicitly authorizes reconnaissance. No other target is acceptable. Use HackerOne — Open programs, Bugcrowd — Public programs or Intigriti. Stay strictly passive: this lab includes no scan and no request to the target.
Step 1 — Pick the target (10 min)
Open the list of public programs. Choose a program that:
- Explicitly authorizes
*.exemple.comin its scope. - Does not require prior registration for passive reconnaissance.
- Concerns a mid-size company (not Google — too massive, and not a 5-person startup — too thin in OSINT). A B2B SaaS with 100 to 1,000 employees is ideal.
Note:
Program chosen: <name>
Platform: HackerOne / Bugcrowd / Intigriti
Program URL: ...
Root domain: target.com
In-scope: *.target.com, target.io
Out of scope: blog.target.com, salesforce.target.com (third-party SaaS)
Disclosure policy: responsible disclosure, 90 days
This block goes at the top of your fiche-cible.md file. It proves you are authorized. It is your RoE for this lab.
Step 2 — Prepare the folder (2 min)
CIBLE=<votre-cible-sans-www> # e.g. acme.io
mkdir -p ~/osint/$CIBLE/{dns,subs,tech,people,leaks,archive}
cd ~/osint/$CIBLE
touch fiche-cible.md
Start fiche-cible.md with the scope block from step 1. Without it, you have no valid excuse if someone asks "why were you looking at my company?".
Step 3 — DNS and email posture (10 min)
Replay the walkthrough loop:
for t in A AAAA MX NS TXT SOA CAA; do
echo "=== $t ==="
dig +short $CIBLE $t
done | tee dns/racine.txt
dig +short _dmarc.$CIBLE TXT | tee dns/dmarc.txt
Answer, in fiche-cible.md, these questions:
- Is the domain behind a CDN? Which one?
- Who are the mail providers?
- Does SPF cite SaaS services? Which ones?
- Is the DMARC policy
none,quarantineorreject? (none = phishing vector)
Step 4 — Subdomains (20 min)
Three sources, to cross-check. Strictly passive.
# Source 1 — Certificate Transparency
curl -s "https://crt.sh/?q=%25.$CIBLE&output=json" \
| jq -r '.[].name_value' \
| tr '[:upper:]' '[:lower:]' \
| sed 's/^\*\.//' \
| sort -u > subs/from-crt.txt
# Source 2 — subfinder
subfinder -d $CIBLE -silent -all -o subs/from-subfinder.txt
# Source 3 — passive amass
amass enum -passive -d $CIBLE -o subs/from-amass.txt
# Merge
cat subs/from-*.txt | sort -u > subs/all-passive.txt
wc -l subs/all-passive.txt
Passive resolution (through public resolvers, never at the client):
dnsx -l subs/all-passive.txt -a -resp -silent -r 1.1.1.1,8.8.8.8 \
-o subs/resolved.txt
Open subs/resolved.txt. Spot:
- IPs outside the CDN (those that do not start with 104.16-31, 172.64-71, 141.101, 173.245 — the Cloudflare blocks — or with those of Fastly, Akamai, and so on). These IPs often point to the real server behind the CDN. Gold target.
- Subdomains with revealing keywords:
dev,staging,test,preprod,internal,old,legacy,admin,panel,jenkins,gitlab,sonar,vault,backup,vpn,ssh,rdp.
- Private IPs (10.x, 172.16-31.x, 192.168.x) published in public DNS: a misconfiguration to report.
In fiche-cible.md, list the 10 most interesting subdomains with their IP and a note on what makes them interesting.
Step 5 — Technologies (20 min)
Without touching the target. Use:
- BuiltWith in your browser.
- Shodan in CLI or on the web:
# For each off-CDN IP identified in step 4:
shodan host <IP> > tech/shodan-<sous-domaine>.txt
- Wappalyzer — watch this, active. If your lab RoE accepts light active, you may use it (a simple page open in the browser = one HTTP hit visible in the target's logs). Otherwise, skip it.
Note, in tech/synthese.md, for each identified target:
- Web server and version.
- Application framework if identifiable.
- Open ports seen by Shodan (not by you — Shodan already scanned them, you consult them).
- Service banners (SSH, mail…).
- Visible misconfigurations: DEBUG mode, error pages that leak the stack, a
Server:header that talks too much.
Step 6 — People (20 min)
# Search emails + LinkedIn profiles (via Bing)
theHarvester -d $CIBLE -b linkedin,bing,duckduckgo -l 500 \
-f people/harvester.html
Complete by hand on LinkedIn:
- Current employees with tech titles (CTO, VP Eng, DevOps, SRE, Sec, IT).
- Former employees (last role:
<your-target>, mentionleft in <date>). Their public GitHubs sometimes hold leftovers.
Consolidate people/personnes.csv:
prenom,nom,titre,email,linkedin,github
Goal: at least 15 people identified. You will attack only 3 or 4 of them later, but you need the pool.
Step 7 — Leaks (20 min)
Three sources, in this order:
1. HaveIBeenPwned by domain (if you have an API key; otherwise use the web UI email by email):
curl -s -H "hibp-api-key: $HIBP_KEY" \
"https://haveibeenpwned.com/api/v3/breacheddomain/$CIBLE" \
| jq . > leaks/hibp.json
2. GitHub Search for the target and for each listed employee:
https://github.com/search?type=commits&q=%22$CIBLE%22
https://github.com/search?type=code&q=%22$CIBLE%22+password
https://github.com/search?type=code&q=%22$CIBLE%22+AKIA
https://github.com/search?type=code&q=%22$CIBLE%22+secret
The AKIA pattern looks for AWS keys. The password pattern looks for… passwords. You will not always find something; when you do, it is gold.
3. Pastebin / Ghostbin / Rentry via Google:
site:pastebin.com "$CIBLE"
site:ghostbin.co "$CIBLE"
"$CIBLE" password ext:txt
Note in leaks/synthese.md:
- Number of target emails leaked and in which dumps.
- Any key, any password, any secret found in the last 3 years.
Step 8 — Wayback archives (10 min)
curl -s "https://web.archive.org/cdx/search/cdx?url=$CIBLE/*&output=json&limit=500" \
| jq -r '.[1:] | .[] | .[2]' \
| sort -u > archive/urls.txt
grep -iE 'admin|backup|debug|swagger|api-doc|test|internal|dev|graphql|env|config|status' archive/urls.txt \
> archive/interessants.txt
wc -l archive/urls.txt archive/interessants.txt
Note, in fiche-cible.md, the 5 to 10 most interesting URLs found in the archives — to revalidate in week 4.
Step 9 — Google dorks (15 min)
Open Google (or DuckDuckGo, quieter) and run at least these queries, with variations:
site:$CIBLE inurl:admin
site:$CIBLE inurl:login
site:$CIBLE inurl:.git
site:$CIBLE ext:sql
site:$CIBLE ext:env
site:$CIBLE ext:log
site:$CIBLE ext:bak
site:$CIBLE intitle:"index of"
site:$CIBLE "internal use only"
site:$CIBLE "confidential"
site:$CIBLE "-----BEGIN" # start of PGP/RSA keys
site:*.amazonaws.com "$CIBLE" # named S3 buckets
site:*.blob.core.windows.net "$CIBLE" # Azure blobs
site:*.digitaloceanspaces.com "$CIBLE"
Screenshot every non-trivial result. Do not click from Google (the click goes through a Google redirector but lands on the target = active). Copy-paste the URL into a curl -I -A "Mozilla/5.0" window later, in week 4, as assumed active work.
Step 10 — Consolidation (30 min)
Open fiche-cible.md and fill the following sections rigorously. Each section has completeness criteria.
# Target sheet — <CIBLE> (<date>)
## 0. Scope and authorization
[the step 1 block, with the program URL]
## 1. Confirmed assets
- Root domain + registrar
- <N> subdomains collected (sources: crt.sh, subfinder, amass)
- Most interesting subdomains (5-10, with IP + note)
## 2. Addressing
- Blocks and providers (Cloudflare, AWS, Azure, GCP, OVH, Vultr…)
- Off-CDN IPs identified
## 3. Technologies
- Front, backend, likely database
- Web servers and versions (via Shodan)
- Related SaaS (Stripe, Twilio, Sentry, Datadog… seen in SPF or headers)
## 4. Email posture
- MX
- SPF (list of includes)
- DMARC (policy, alignment)
- DKIM present?
## 5. Key people (>= 15)
- people/personnes.csv file attached
## 6. Public leaks
- HIBP: <N> accounts involved, dumps
- GitHub: secrets found (specify repo, file, secret type)
- Pastebin / Google: finds
## 7. Five priority targets (descending order)
1. <target> — <reason> — <estimated effort to confirm>
2. ...
3. ...
4. ...
5. ...
## 8. Next step (module 4)
- Active-phase plan: which subdomain, which tool, which window.
- Points to negotiate in the RoE (if a subdomain was found outside the original scope).
Self-assessment grid
- The bug bounty program and its scope are cited at the top of the sheet.
- At least 50 subdomains collected (aim for 100+ on a normal target).
- Cross-check done:
crt.sh+subfinder+amass(all 3, not just one). - Every IP is classified CDN / non-CDN / private.
- Email posture documented (SPF, DMARC, DKIM).
- At least 15 people with their email guessed from the pattern.
- At least one leak source consulted (HIBP or GitHub Search).
- Wayback Machine consulted, top URLs noted.
- Five priority targets ranked with justification.
- No command touched the target. No scan. No direct
curl.
All ticked? Your sheet is ready to feed week 4.
What usually blocks you
- Few subdomains: the target is small, or mostly behind a CDN that hides well. Spend more time on GitHub and LinkedIn.
- No access to Shodan/HIBP: free tiers exist, thinner. Or wait until the end of the course — the subscription is worth it.
- theHarvester returns little: LinkedIn regularly shuts Bing access. Build the org chart by hand; it is 30 minutes.
- The urge to scan: that is the sign you have not dug the passive sources far enough. Go back and search.
What you take away from this lab
- A reproducible workflow in 2 h on any authorized target.
- A 2-page sheet that will be the reference for every following module on this target.
- The instinct to recognize an interesting subdomain in one second.
- The discipline to send nothing before you have looked at everything.