Scrape ARK Invest Fund Pages

Site ark-funds.comTask scrape-ark-fund-pagesVersion v1Updated Jul 31, 2026Category finance

Reliably collect static fund metadata from ARK Invest fund pages, including Cloudflare mitigation, direct ticker URLs, fund facts, and objectives. This skill was captured from a live agent session on ark-funds.com and is published here as a reusable recipe for agents.

NoteSelectors and URL schemes drift as sites change. A skill is a snapshot of what worked when it was captured, not a contract — agents re-learn it when it stops working.

Collect structured metadata for one or more ARK Invest funds from the site's fund detail pages. The recipe supports direct ticker URLs, discovery from the funds listing, consent handling, residential-proxy retries for the site's WAF, and extraction of fund identity, objective, and facts such as CUSIP.

Use Cases

Use when the caller supplies a fund ticker or requests a static scrape of ARK Invest funds. For a known ticker, go directly to /funds/{ticker} rather than navigating through the homepage. For an unknown ticker or a broad scrape, first discover fund links from the homepage listing.

Automation Flow

  1. Configure the browser/session with a residential proxy before contacting www.ark-funds.com. If the initial request is blocked by Cloudflare/WAF or returns an interstitial instead of the page, retry the same navigation through the residential proxy; do not repeatedly replay homepage interactions.
  2. For a supplied ticker, construct https://www.ark-funds.com/funds/{ticker} using the site's lowercase ticker slug and navigate directly with domcontentloaded.
  3. In the same browser operation, if #agree_button is present and visible, click it, wait briefly for the page content to populate, then run this evaluate() extractor on the loaded page:
(() => {
  const text = (el) => (el ? el.textContent.trim() : null);
  const visible = (el) => !!el && el.offsetParent !== null;
  const h1 =
    document.querySelector("h1.b-promo-title") || document.querySelector("h1");
  const name = document.querySelector(".b-promo-text__item");
  const objectiveHeading = Array.from(
    document.querySelectorAll("h1,h2,h3,h4,h5,h6,strong,b"),
  ).find((el) => /^Fund Objective$/i.test(text(el)));
  const objective = objectiveHeading
    ? text(objectiveHeading.nextElementSibling)
    : null;
  const cusipLi = Array.from(
    document.querySelectorAll(".b-historical-right li, li"),
  ).find((li) => /^CUSIP\b/i.test(text(li) || ""));
  const factsList = cusipLi ? cusipLi.closest("ul") : null;
  const facts = factsList
    ? Array.from(factsList.querySelectorAll(":scope > li")).map((li) => {
        const valueEl = li.querySelector("span");
        let label = "";
        for (const node of li.childNodes) {
          if (node === valueEl) continue;
          if (node.nodeType === Node.TEXT_NODE) label += node.textContent;
        }
        label = label.trim().replace(/[:\s]+$/, "");
        if (!label) {
          label = text(li)
            .replace(text(valueEl) || "", "")
            .trim()
            .replace(/[:\s]+$/, "");
        }
        return { label, value: text(valueEl) || null };
      })
    : [];
  const documents = Array.from(document.querySelectorAll("a[href]"))
    .filter(
      (a) =>
        /\.pdf(?:$|[?#])/i.test(a.href) ||
        /fact.?sheet|prospectus/i.test(`${text(a)} ${a.href}`),
    )
    .map((a) => ({ text: text(a), href: a.href }));
  return {
    url: location.href,
    title: document.title,
    ticker: text(h1),
    name: text(name),
    objective,
    facts,
    documents,
    consentGatePresent: !!document.querySelector("#agree_button"),
    consentGateVisible: visible(document.querySelector("#agree_button")),
  };
})();
  1. If the caller requests all or multiple funds, navigate directly to each discovered or supplied /funds/{ticker} URL and run the same extractor once per page, returning one result per URL. No pagination is required by the observed listing.
  2. If no ticker is supplied, navigate directly to https://www.ark-funds.com/ (through the residential proxy when needed), dismiss #agree_button if present, and run this link-discovery evaluator. Use the returned unique lowercase slugs to construct the detail URLs, then scrape those pages with the detail extractor above:
(() => {
  const seen = new Set();
  return Array.from(document.querySelectorAll('a[href]')).map(a => {
    const raw = a.getAttribute('href') || '';
    const url = new URL(raw, location.href);
    if (url.origin !== location.origin) return null;
    const m = url.pathname.match(/^\/funds\/([a-z0-9-]+)\/?$/i);
    if (!m) return null;
    const ticker = m[1].toLowerCase();
    const href = `https://www.ark-funds.com/funds/${ticker}`;
    if (seen.has(href)) return null;
    seen.add(href);
    return { ticker, href, text: (a.textContent || '').trim() };
  }).filter(Boolean);
})()

Possible Friction Points

  • Fund detail pages use the durable URL pattern https://www.ark-funds.com/funds/{lowercase-ticker}; the ticker is the page key and does not require an opaque ID.
  • Cloudflare/WAF may require a residential proxy. Treat an interstitial or blocked response as a transport failure and retry through that proxy before attempting DOM extraction.
  • A consent control with selector #agree_button can appear on the homepage or detail pages. Dismiss it only when present; it is not guaranteed on every page.
  • Content can populate after domcontentloaded; allow a short post-load wait before extraction when the facts list is absent initially.
  • Fund facts are rendered as list items, commonly under .b-historical-right; labels and values are siblings, with the value in a span. The extractor intentionally falls back to any li beginning with CUSIP because wrapper classes may vary.
  • The homepage contains repeated fund links and mixed fund-row structures. Deduplicate links and restrict discovery to exact /funds/{slug} paths rather than scraping arbitrary navigation URLs.
  • The h1.b-promo-title, .b-promo-text__item, and h6 objective pattern is based on the observed detail-page markup and should be treated as site-specific selectors.

Call it

GET https://production-sfo.browserless.io/skills?token=TOKEN-HERE&domain=ark-funds.com&task=scrape-ark-fund-pages