All modules

Module 03 · Operator

Reconnaissance & OSINT

Learn passive and active reconnaissance: OSINT, subdomain discovery, certificate transparency, and building an attack surface inventory without rushing to exploits.

2 lessons
11 deep sections
~115 min guided
Progress…

Authorized testing only. Practice on systems you own, isolated labs, or targets with written permission. Unauthorized access is illegal.

Outcomes

  • Build a passive footprint for a domain you are authorized to test
  • Use modern recon pipelines (subfinder → resolve → httpx) responsibly
  • Organize assets into a testable, scoped inventory
  • Know when recon is “done enough” to begin scanning
  • Document sources and timestamps so recon is reproducible and report-ready

Lessons

Lesson 1
55 min5 sections

Passive recon first

Collect without touching the target: CT logs, search operators, archives, WHOIS, and internet-wide search engines — under authorization and policy.

Learning objectives

  • Explain passive vs active recon and when each is appropriate
  • Use certificate transparency, WHOIS, and search operators for inventory
  • Harvest historical URLs and parameters from archives
  • Document public exposures responsibly without over-collection of PII
  • Turn raw OSINT into a scoped asset list with confidence levels

Deep teach-through

Passive vs active: noise, ethics, and signal

Passive recon gathers data from third parties and public sources without sending packets directly to the target’s infrastructure. Active recon probes the target (DNS queries to their servers, HTTP requests, port scans). In practice the line can blur (e.g., DNS to authoritative servers), so judge by intent, policy, and what the RoE allows.

Start passive when you need breadth with less noise: forgotten subdomains, old staging names, leaked documents, employee naming patterns, technology hints. Move active when passive inventory is organized and you have written authorization for the addresses and names you will touch.

Authorization still applies. Querying public CT logs for a domain in a bug bounty program may be fine; bulk scraping personal data, harassing individuals, or using breach datasets against people without a lawful professional purpose is not “just OSINT.” Stay mission-focused and minimize sensitive personal data in notes.

Certificate Transparency and DNS-oriented inventory

Certificate Transparency (crt.sh and similar) logs certificates issued for names under a domain. Subdomains appear because operators requested certs — including staging, VPN, and legacy systems. Query, deduplicate, and mark which names still resolve.

WHOIS and RDAP give registration metadata, name servers, and sometimes contacts. Privacy redaction is common; still record registrar and NS for attack-surface and phishing-resistance discussions. dig/nslookup against public resolvers maps what the world sees for A/AAAA/MX/TXT/CNAME.

Subdomain enumeration tools (subfinder, Amass in passive modes, theHarvester) automate multi-source collection. Prefer passive/source modes first; understand each tool’s configuration so you do not accidentally enable aggressive active modes against out-of-scope infrastructure.

Search operators, archives, and code leaks

Search engines still reward skilled operators: site:, filetype:, inurl:, and careful phrases surface exposed indexes, docs, and login portals. Google dorking is not magic — it is disciplined querying plus note-taking. Respect robots and legal boundaries; do not automate abusive scraping of search engines.

Web archives (Wayback and similar) preserve old paths, parameters, and JS endpoints that modern SPAs still honor. waybackurls and related tooling help extract historical URLs for later content discovery. Filter for admin, api, backup, .git, and auth-related paths — always against in-scope hosts only when you move to active checks.

Public code hosts, paste sites, and cloud buckets sometimes leak secrets. If you find live credentials during authorized testing, handle as sensitive loot: notify per RoE, do not reuse outside the engagement, and never dump findings on social media for clout.

Internet-wide search engines and threat intel views

Shodan, Censys, and similar engines show banners and services already observed by third parties. They are excellent for scope discussions (“this management port is already public”) and for prioritizing hardening. Using them to attack random hosts remains unauthorized activity.

Interpret results with humility: banners lag, CDN front-ends confuse ownership, and shared hosting may show neighbors. Cross-check with DNS and the client’s asset list before treating an IP as in scope.

Record the query, date, and screenshot or export. Recon claims in reports should be reproducible. If a port was open on Shodan but closed during your test, that historical exposure can still be a useful narrative for risk and monitoring gaps.

From piles of data to an attack-surface inventory

Raw tool output is not inventory. Normalize into fields: name, IP(s), source (CT, search, client list), resolves? (yes/no), in-scope? (yes/no/unknown), notes. Unknown scope items become questions for the client, not silent additions to your scan list.

Assign confidence and priority: production marketing site vs forgotten Jenkins on a weird subdomain. Passive recon often finds the weird ones — those drive high-value paths later.

Stop conditions matter. Infinite OSINT is a real failure mode. Define “done enough for this phase”: core domains covered, major sources queried, inventory reviewed, open questions listed. Then proceed to pipelines and scanning under RoE.

Key concepts

Passive recon
Collection from public/third-party sources without direct interaction with the target systems.
Certificate Transparency
Public logs of issued TLS certificates that often reveal hostnames and subdomains.
Attack surface inventory
Normalized list of names, addresses, and services in scope with sources and status.
Historical URL
Path or endpoint observed in archives or old content that may still exist on live apps.
Scope confidence
Your assessed certainty that an asset is authorized — unknown means ask, do not assume.

Common mistakes

  • Jumping to port scans before organizing passive inventory
  • Adding “interesting” subdomains to active tests without scope confirmation
  • Hoarding personal data unrelated to the engagement objective
  • Treating Shodan hits as automatic permission to exploit
  • No timestamps or source tags, making recon unreproducible
  • Running active Amass/bruteforce modes thinking they are passive

Defender view

  • CT logs and public banners are what adversaries see first — monitor your own exposure.
  • Stale DNS and forgotten subdomains are recurring breach entry points.
  • Reducing public tech disclosure and locking down staging naming helps everyone.

Operator checklist

  • I can point to the policy or RoE covering this domain/org
  • Passive sources queried are listed with dates in notes
  • Inventory distinguishes in-scope, out-of-scope, and unknown
  • PII and secrets are minimized and stored per data-handling rules
  • I have a written stop condition for this recon phase

Example commands & patterns

# Only against domains you own or are authorized to assess
curl -s 'https://crt.sh/?q=%25.example.com&output=json' | jq -r '.[].name_value' | sort -u
whois example.com
dig example.com NS +short
subfinder -d example.com -silent -o subs-passive.txt
echo example.com | waybackurls | sort -u | tee wayback.txt
# Shodan CLI if configured: shodan domain example.com

Practice drills

  1. Pick a domain you own (or a public training target with clear permission) and list 20+ subdomains passively
  2. Pull wayback URLs and filter for login, admin, api, and backup-related paths
  3. Run WHOIS/RDAP and dig; document registrar, NS, and notable TXT records
  4. Build a markdown inventory table with columns: asset, source, resolves, scope status
  5. Write three client questions raised by “unknown scope” assets you found

Next: Chain discovery into a disciplined resolve → probe pipeline without thrashing targets.

Lesson 2
60 min6 sections

Modern recon pipelines

Chain tools like operators: discover → resolve → probe → light tech checks — with scope files, rate limits, and durable storage.

Learning objectives

  • Build a shell pipeline with ProjectDiscovery-style tools under authorization
  • Deduplicate and store results with timestamps and tool versions
  • Apply rate limits and avoid thrashing production targets
  • Produce clean live-host and HTTP title inventories for the next phase
  • Know when to stop automating and start manual analysis

Deep teach-through

Pipeline thinking: stages and handoffs

A modern recon pipeline is a staged factory: (1) discover names, (2) resolve to IPs, (3) optional port foreshadowing, (4) HTTP probe for live web, (5) light fingerprint or safe template checks, (6) manual triage. Each stage has inputs, outputs, and failure modes.

Example flow (authorized targets only): subfinder → resolve (dnsx/puredns/massdns style tooling as available) → httpx → optional katana crawl on approved hosts → selective nuclei with safe tags → notes. Adjust for RoE: some programs forbid automated vulnerability scanning entirely.

Never let a wildcard mental model send traffic out of scope. Physical files win: scope.txt, oos.txt, roots.txt. Feed tools explicit lists. If a new name appears mid-pipeline, gate it through scope review before active stages.

Discovery and resolution hygiene

Combine passive sources first, then consider permitted active DNS brute force only if RoE allows and rate is sane. Wordlists should match the environment (corp naming vs product SaaS). Longer lists are not automatically better if they burn time and trigger defenses.

Resolution must handle wildcards. Some domains resolve every label; without wildcard detection you invent infinite fake hosts. Record both names and IPs; CDNs may map many names to shared edges — ownership still follows the name and client asset rules.

Deduplicate with sort -u and tools like anew so daily runs only add deltas. Recon is iterative across days of an engagement; deltas are where new staging systems appear.

HTTP probing and crawl discipline

httpx (and similar) confirms which hosts speak HTTP(S), captures status, title, tech hints, and TLS info. Tune threads and rate. Production marketing sites and fragile appliances are not CTF boxes — start slow, increase only if stable.

Crawlers (katana, browser-based tools) expand paths on live apps. Stay on in-scope hosts, respect authentication boundaries, and avoid destructive methods. Save URL lists for content discovery later; do not weaponize every path immediately.

Titles and status codes feed prioritization: “Jenkins”, “login”, “phpmyadmin”, default framework pages, and error banners often beat generic marketing sites for early attention — after scope confirmation.

Safe automation and nuclei temptation

Nuclei and similar template engines are powerful and easy to misconfigure. Use only with authorization, against in-scope assets, with templates appropriate to the engagement (often “safe” or informational first). Aggressive CVE spray can DoS or violate bounty rules.

Automation without notes is noise. For every automated alert you pursue, capture the request/response evidence and a human hypothesis. False positives are common; your reputation depends on verification.

naabu and other port foreshadowing tools may sit between resolve and full Nmap. Keep methodology staged: do not confuse a quick top-ports peek with full enumeration (Module 4). Document what each stage was meant to answer.

Storage, JSON, and collaboration

Save raw outputs with timestamps: scans/recon/2026-08-10_subfinder.txt, httpx.json, etc. Prefer machine-readable exports when available and use jq to build clean lists (live-hosts.txt, titles.txt, techs.txt).

A teammate should be able to re-run or continue from your folder. Include tool versions (subfinder -version) in notes. Pipelines evolve; reproducibility prevents “it found it yesterday” mystery.

Know when to stop automating. If httpx shows three interesting apps, deep manual mapping may beat another hour of template spam. Experts schedule thinking time; beginners hide in tools.

Rate limits, legality, and professionalism

Even authorized tests can harm availability. Use rate flags, avoid business-critical windows if RoE says so, and coordinate with the client when automation must be louder. Bug bounties often ban automated scanning or limit it — read the policy again before pipelines.

Third-party SaaS in the inventory may be out of scope for active probing even if DNS points at them. Separate “client-controlled” from “vendor-controlled” assets.

Your pipeline is part of your professional brand. Clean inventories, low collateral noise, and clear handoff to scanning/exploitation phases mark operator maturity more than a huge raw subdomain count.

Key concepts

Recon pipeline
Staged, scriptable workflow moving from names to live services with explicit handoff artifacts.
Wildcard DNS
Configuration that resolves arbitrary sublabels, creating false inventory if undetected.
Delta recon
Re-running collection to capture only new assets since the last run.
HTTP probe
Lightweight request to determine liveness, status, titles, and basic tech signals.
Safe templates
Low-impact automated checks that avoid exploitation and availability harm when policy allows.

Common mistakes

  • Piping discovery straight into aggressive scanners with no scope file
  • Ignoring wildcard DNS and storing thousands of fake hosts
  • Maxing threads against production “because the blog did”
  • Running full nuclei CVE packs on bounty targets that forbid it
  • Losing raw outputs and only keeping a screenshot of a terminal
  • Never converting JSON outputs into clean lists for the next stage

Defender view

  • Consistent recon noise from one source IP range is easy to detect and block — expect controls.
  • Asset management that matches CT and DNS reality shrinks surprise subdomains.
  • Rate limiting and bot management protect origin apps from careless automation.

Operator checklist

  • scope.txt / roots.txt drive every active stage
  • Discovery output is deduplicated and timestamped
  • Wildcard checks done before trusting resolved lists
  • httpx/crawl rates are conservative until stability is clear
  • live-hosts.txt and titles.txt exist for prioritization

Example commands & patterns

# Authorized lab or owned domain only
subfinder -dL roots.txt -silent | anew subs.txt
cat subs.txt | dnsx -silent -a -resp | tee resolved.txt
cat subs.txt | httpx -silent -title -tech-detect -status-code -o httpx.txt
cat httpx.txt | anew live-http.txt
katana -list live-http.txt -d 2 -o urls.txt
# Only if RoE allows automated checks:
nuclei -l live-http.txt -tags tech,exposure -rate-limit 20 -o nuclei-safe.txt
jq -r '.' httpx.jsonl 2>/dev/null | head

Practice drills

  1. Run a full passive→resolve→httpx pipeline against your lab domain or a clearly permitted demo target
  2. Produce clean live-hosts.txt and titles.txt and attach them to engagement notes
  3. Demonstrate anew by running discovery twice and showing only new lines appended
  4. Intentionally lower rate limits and document the flags you used for a “production-safe” profile
  5. Triage five httpx titles into priority order with written reasons (not just “looks cool”)

Next: With a live inventory, move to Module 4 for staged scanning and service enumeration — still under authorization.