[{"data":1,"prerenderedAt":2669},["ShallowReactive",2],{"blog-guide-\u002Fblog\u002Fguides\u002Fingest-mixed-documents-llm-embedding":3,"blog-guide-related-\u002Fblog\u002Fguides\u002Fingest-mixed-documents-llm-embedding":1086},{"id":4,"title":5,"author":6,"body":7,"category":1064,"cover":1065,"description":1066,"draft":1067,"extension":1068,"image":1069,"launchCta":1069,"listingCover":1070,"meta":1071,"navigation":1072,"ogImage":1069,"path":1073,"publishedAt":1074,"readTime":1075,"seo":1076,"stem":1077,"tags":1078,"toolCategory":936,"updatedAt":1084,"__hash__":1085},"blogGuides\u002Fblog\u002Fguides\u002Fingest-mixed-documents-llm-embedding.md","Ingesting Mixed Documents for LLM Embedding: PDF to Markdown","Jasper Li",{"type":8,"value":9,"toc":1037},"minimark",[10,15,78,87,92,95,100,103,106,110,113,116,120,128,132,210,213,228,232,235,239,250,258,267,271,328,332,338,350,355,414,436,443,453,457,462,482,486,520,529,537,541,546,549,552,556,570,574,577,581,609,642,653,657,667,676,679,683,686,694,698,704,708,711,714,717,721,908,915,919,922,925,928,933,941,945,948,951,954,957,961,964,967,983,987,1002,1008,1014,1027,1033],[11,12,14],"p",{"style":13},"font-size:18px !important;line-height:1.65 !important;margin:0 0 24px;color:inherit;","Copy this line to your agent to turn a folder of mixed files into clean Markdown.",[16,17,22],"pre",{"className":18,"code":19,"language":20,"meta":21,"style":21},"language-sh shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","set up https:\u002F\u002Fmonid.ai\u002FSKILL.md and use context.dev \u002Fparse to convert these PDFs and Office files to Markdown\n","sh","",[23,24,25],"code",{"__ignoreMap":21},[26,27,30,34,38,41,44,47,50,53,56,59,62,65,67,70,73,75],"span",{"class":28,"line":29},"line",1,[26,31,33],{"class":32},"s2Zo4","set",[26,35,37],{"class":36},"sfazB"," up",[26,39,40],{"class":36}," https:\u002F\u002Fmonid.ai\u002FSKILL.md",[26,42,43],{"class":36}," and",[26,45,46],{"class":36}," use",[26,48,49],{"class":36}," context.dev",[26,51,52],{"class":36}," \u002Fparse",[26,54,55],{"class":36}," to",[26,57,58],{"class":36}," convert",[26,60,61],{"class":36}," these",[26,63,64],{"class":36}," PDFs",[26,66,43],{"class":36},[26,68,69],{"class":36}," Office",[26,71,72],{"class":36}," files",[26,74,55],{"class":36},[26,76,77],{"class":36}," Markdown\n",[11,79,80,81,86],{},"Every retrieval pipeline has the same first problem and almost nobody writes it down: the documents are not one format. A knowledge base is PDFs, DOCX, a few spreadsheets, some scanned faxes and a pile of web pages, and the embedding model wants one clean text stream. The conversion step is where quality is won or lost, well before chunking or retrieval, and it is the step that gets three lines of code and no tests. Monid is ",[82,83,85],"a",{"href":84},"\u002Fopenrouter-for-agent-tools","the OpenRouter for agent tools",": one key and one balance across many providers, so the conversion step can be one endpoint instead of six libraries.",[88,89,91],"h2",{"id":90},"why-is-ingesting-mixed-document-types-harder-than-it-looks","Why is ingesting mixed document types harder than it looks?",[11,93,94],{},"Because format conversion is not one problem, it is four, and they fail at different points. The r\u002FLocalLLaMA thread that names this best lists them in the title: PDF, Office and HTML conversion, OCR, de-duplication and chunking. Twelve comments in, the consensus is that people underestimate the first and over-engineer the last.",[96,97,99],"h3",{"id":98},"a-pdf-is-a-layout-format-not-a-document-format","A PDF is a layout format, not a document format",[11,101,102],{},"PDF describes where marks go on a page. It does not describe reading order, and a two-column academic paper or a table-heavy report will extract into interleaved nonsense if the extractor reads coordinates naively. The text is all there and it is in the wrong order, which is worse than missing text, because it embeds cleanly and retrieves garbage.",[11,104,105],{},"This is the failure that survives all the way to production, because nothing errors. A chunk of interleaved column text is a valid chunk with a valid embedding. It just answers no question correctly.",[96,107,109],{"id":108},"scanned-documents-have-no-text-at-all","Scanned documents have no text at all",[11,111,112],{},"A scanned contract or a fax is an image inside a PDF wrapper. A text extractor returns an empty string and reports success. If your ingestion pipeline logs a page count rather than a character count, an entire scanned archive can pass through it and produce nothing, silently.",[11,114,115],{},"The check that catches this is one line: assert a minimum character count per page, and route anything below it to OCR rather than to the embedder.",[96,117,119],{"id":118},"html-brings-furniture-you-did-not-ask-for","HTML brings furniture you did not ask for",[11,121,122,123,127],{},"Web pages carry navigation, footers, cookie banners and sidebars, and all of it embeds. We measured the extreme version of this in ",[82,124,126],{"href":125},"\u002Fblog\u002Fguides\u002Ffree-api-extract-page-content-rag","a free API to extract page content for RAG",": a Wikipedia page came back at 74,552 characters through a naive scrape, starting with the nav menu, against 689 clean ones through the structured route. Those extra characters are not merely wasted tokens, they dilute the embedding of every chunk they land in.",[96,129,131],{"id":130},"one-library-per-format-versus-one-conversion-endpoint","One library per format versus one conversion endpoint",[133,134,135,151],"table",{},[136,137,138],"thead",{},[139,140,141,145,148],"tr",{},[142,143,144],"th",{},"Aspect",[142,146,147],{},"A library per format",[142,149,150],{},"One conversion endpoint",[152,153,154,166,177,188,199],"tbody",{},[139,155,156,160,163],{},[157,158,159],"td",{},"Setup",[157,161,162],{},"Six dependencies, native builds for OCR",[157,164,165],{},"One HTTP call",[139,167,168,171,174],{},[157,169,170],{},"Coverage gaps",[157,172,173],{},"Found in production, one format at a time",[157,175,176],{},"Format list is published up front",[139,178,179,182,185],{},[157,180,181],{},"OCR",[157,183,184],{},"Separate install, separate tuning",[157,186,187],{},"A boolean on the same call",[139,189,190,193,196],{},[157,191,192],{},"Failure mode",[157,194,195],{},"Silent empty string per format",[157,197,198],{},"One response shape to assert on",[139,200,201,204,207],{},[157,202,203],{},"Best for",[157,205,206],{},"Full control and offline processing",[157,208,209],{},"Getting the corpus in this week",[11,211,212],{},"The pattern is control against surface area. Local libraries are the right answer when documents cannot leave your network, and the wrong answer when the real cost is six half-maintained format handlers.",[214,215,216],"blockquote",{},[11,217,218,219,223,224],{},"📖 ",[220,221,222],"strong",{},"See also"," ",[82,225,227],{"href":226},"\u002Fblog\u002Fany-url-to-llm-ready-markdown","Any URL to LLM-Ready Markdown: A Copy-Paste Cookbook",[88,229,231],{"id":230},"what-are-the-best-practices-for-ingesting-mixed-document-types-for-llm-extraction","What are the best practices for ingesting mixed document types for LLM extraction?",[11,233,234],{},"Convert first, normalise second, chunk last, and put the assertions between the steps rather than at the end. The order matters because each step can fail quietly, and a check after conversion costs nothing while a check after embedding costs a reindex.",[96,236,238],{"id":237},"for-agents","For agents",[11,240,241,242,249],{},"Grab an API key at ",[82,243,248],{"href":244,"rel":245},"https:\u002F\u002Fapp.monid.ai\u002F",[246,247],"noopener","noreferrer","app.monid.ai",", then paste this to your agent and hand it the key:",[16,251,256],{"className":252,"code":254,"language":255,"meta":21},[253],"language-text","set up https:\u002F\u002Fmonid.ai\u002FSKILL.md\n","text",[23,257,254],{"__ignoreMap":21},[11,259,260,261,266],{},"It learns the whole discover, inspect, run workflow itself. More in the ",[82,262,265],{"href":263,"rel":264},"https:\u002F\u002Fmonid.ai\u002Fdocs\u002Fguide\u002Fquickstart-skill",[246,247],"agent quickstart",".",[96,268,270],{"id":269},"for-humans","For humans",[16,272,276],{"className":273,"code":274,"language":275,"meta":21,"style":21},"language-bash shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","npm install -g @monid-ai\u002Fcli\nmonid keys add -k \u003Cyour-key> -l main\n","bash",[23,277,278,293],{"__ignoreMap":21},[26,279,280,284,287,290],{"class":28,"line":29},[26,281,283],{"class":282},"sBMFI","npm",[26,285,286],{"class":36}," install",[26,288,289],{"class":36}," -g",[26,291,292],{"class":36}," @monid-ai\u002Fcli\n",[26,294,296,299,302,305,308,312,315,319,322,325],{"class":28,"line":295},2,[26,297,298],{"class":282},"monid",[26,300,301],{"class":36}," keys",[26,303,304],{"class":36}," add",[26,306,307],{"class":36}," -k",[26,309,311],{"class":310},"sMK4o"," \u003C",[26,313,314],{"class":36},"your-ke",[26,316,318],{"class":317},"sTEyZ","y",[26,320,321],{"class":310},">",[26,323,324],{"class":36}," -l",[26,326,327],{"class":36}," main\n",[96,329,331],{"id":330},"step-1-convert-every-format-through-one-door","Step 1. Convert every format through one door",[11,333,334,337],{},[220,335,336],{},"What it does."," Takes a file at a URL and returns GitHub Flavored Markdown, across more than sixty formats, so the branch in your code is data rather than control flow.",[11,339,340,223,343,349],{},[220,341,342],{},"The endpoints.",[82,344,346],{"href":345},"\u002Ftools\u002Fweb",[23,347,348],{},"context.dev\u002Fparse"," handles PDF, DOCX, XLSX, PPTX, RTF, HTML, images, code files and structured data including JSON, CSV, YAML and XML, from any public HTTPS URL up to 25MB.",[11,351,352],{},[220,353,354],{},"The call.",[16,356,358],{"className":273,"code":357,"language":275,"meta":21,"style":21},"monid inspect -p context.dev -e \u002Fparse\nmonid run -p context.dev -e \u002Fparse \\\n  -i '{\"file_url\":\"https:\u002F\u002Fexample.com\u002Freport.pdf\",\"useMainContentOnly\":true}' -w\n",[23,359,360,378,396],{"__ignoreMap":21},[26,361,362,364,367,370,372,375],{"class":28,"line":29},[26,363,298],{"class":282},[26,365,366],{"class":36}," inspect",[26,368,369],{"class":36}," -p",[26,371,49],{"class":36},[26,373,374],{"class":36}," -e",[26,376,377],{"class":36}," \u002Fparse\n",[26,379,380,382,385,387,389,391,393],{"class":28,"line":295},[26,381,298],{"class":282},[26,383,384],{"class":36}," run",[26,386,369],{"class":36},[26,388,49],{"class":36},[26,390,374],{"class":36},[26,392,52],{"class":36},[26,394,395],{"class":317}," \\\n",[26,397,399,402,405,408,411],{"class":28,"line":398},3,[26,400,401],{"class":36},"  -i",[26,403,404],{"class":310}," '",[26,406,407],{"class":36},"{\"file_url\":\"https:\u002F\u002Fexample.com\u002Freport.pdf\",\"useMainContentOnly\":true}",[26,409,410],{"class":310},"'",[26,412,413],{"class":36}," -w\n",[11,415,416,419,420,423,424,427,428,431,432,435],{},[220,417,418],{},"What comes back."," The parsed document as Markdown. Four switches shape it, verified 2026-08-20: ",[23,421,422],{},"includeLinks"," preserves hyperlinks and defaults on, ",[23,425,426],{},"includeImages"," adds image references and defaults off, ",[23,429,430],{},"shortenBase64Images"," truncates inline image payloads and defaults on, and ",[23,433,434],{},"useMainContentOnly"," drops headers, footers, sidebars and navigation where they can be detected. That last one is the HTML furniture problem solved with a boolean.",[11,437,438,439,442],{},"There is also an ",[23,440,441],{},"extension"," hint for when neither the URL nor the Content-Type reveals the format, which is the case more often than you would like with files pulled from object storage.",[11,444,445,448,449,266],{},[220,446,447],{},"What it costs."," A fraction of a cent per call, billed per call rather than per page, so a two hundred page report costs the same as a one pager. Current figures at ",[82,450,452],{"href":451},"\u002Ftools","monid.ai\u002Ftools",[96,454,456],{"id":455},"step-2-route-the-scanned-files-to-ocr-and-only-those","Step 2. Route the scanned files to OCR, and only those",[11,458,459,461],{},[220,460,336],{}," Detects and reads text from images embedded in PDF pages, which is the only way to get anything at all from a scanned archive.",[11,463,464,466,467,471,472,475,476,481],{},[220,465,342],{}," The same ",[82,468,469],{"href":345},[23,470,348],{}," call with ",[23,473,474],{},"ocr: true",". For loose images rather than PDFs, ",[82,477,478],{"href":345},[23,479,480],{},"api.strale.io\u002Fx402\u002Fimage-to-text"," runs OCR through vision and returns text with a confidence score, which is useful when you need to threshold on quality.",[11,483,484],{},[220,485,354],{},[16,487,489],{"className":273,"code":488,"language":275,"meta":21,"style":21},"monid run -p context.dev -e \u002Fparse \\\n  -i '{\"file_url\":\"https:\u002F\u002Fexample.com\u002Fscanned-contract.pdf\",\"ocr\":true}' -w\n",[23,490,491,507],{"__ignoreMap":21},[26,492,493,495,497,499,501,503,505],{"class":28,"line":29},[26,494,298],{"class":282},[26,496,384],{"class":36},[26,498,369],{"class":36},[26,500,49],{"class":36},[26,502,374],{"class":36},[26,504,52],{"class":36},[26,506,395],{"class":317},[26,508,509,511,513,516,518],{"class":28,"line":295},[26,510,401],{"class":36},[26,512,404],{"class":310},[26,514,515],{"class":36},"{\"file_url\":\"https:\u002F\u002Fexample.com\u002Fscanned-contract.pdf\",\"ocr\":true}",[26,517,410],{"class":310},[26,519,413],{"class":36},[11,521,522,524,525,528],{},[220,523,418],{}," The Markdown, plus an ",[23,526,527],{},"ocr_ran"," field telling you whether OCR actually executed. That field is worth reading rather than ignoring, because it is also how the billing settles.",[11,530,531,533,534,536],{},[220,532,447],{}," This is the one genuinely clever piece of billing in the ingestion path. Setting ",[23,535,474],{}," holds the higher amount, but the OCR line only settles if OCR actually ran, detected from the vendor's own meter. A file with nothing to OCR bills the base rate. So you can set the flag across a mixed batch without paying the OCR price for the text-layer files, which means you do not need to pre-classify the batch yourself.",[96,538,540],{"id":539},"step-3-assert-then-chunk","Step 3. Assert, then chunk",[11,542,543,545],{},[220,544,336],{}," Catches the silent failures before they become embeddings.",[11,547,548],{},"Three assertions cover almost everything that goes wrong. Check a minimum character count per source page and route failures to OCR. Check that the Markdown contains at least one heading if the source had structure, because a document that converted to one undifferentiated block usually lost its reading order. And hash the normalised text before embedding, because de-duplication is far cheaper on strings than on vectors.",[11,550,551],{},"Chunk after all of that. Chunking a bad conversion produces bad chunks efficiently.",[553,554],"skill-prompt",{"prompt":555},"convert these mixed PDFs and Office files to Markdown, turn OCR on, and tell me which files came back with fewer than 200 characters",[214,557,558],{},[11,559,218,560,223,562,565,566],{},[220,561,222],{},[82,563,564],{"href":125},"A Free API to Extract Page Content for RAG"," and ",[82,567,569],{"href":568},"\u002Fblog\u002Fbest-ocr-api-for-messy-images-2026","The Best OCR API for Messy Images in 2026",[88,571,573],{"id":572},"how-do-you-convert-a-pdf-to-markdown","How do you convert a PDF to Markdown?",[11,575,576],{},"One call, and the interesting parts are what happens when the PDF has no text layer and where the output actually lives.",[96,578,580],{"id":579},"the-straightforward-case","The straightforward case",[11,582,583,584,589,590,593,594,597,598,601,602,605,606,266],{},"A born-digital PDF, one where the text is real text rather than a picture of text, converts directly. Running ",[82,585,587],{"href":586},"\u002Ftools\u002Fweb-extraction",[23,588,348],{}," against a public research paper on 2026-08-26 returned ",[23,591,592],{},"success: true",", ",[23,595,596],{},"type: \"pdf\"",", and a ",[23,599,600],{},"document"," object holding a signed ",[23,603,604],{},"download_link"," with ",[23,607,608],{},"content_type: \"text\u002Fmarkdown\"",[16,610,612],{"className":273,"code":611,"language":275,"meta":21,"style":21},"monid run -p context.dev -e \u002Fparse -w -i '{\"file_url\": \"https:\u002F\u002Fexample.com\u002Freport.pdf\"}'\n",[23,613,614],{"__ignoreMap":21},[26,615,616,618,620,622,624,626,628,631,634,636,639],{"class":28,"line":29},[26,617,298],{"class":282},[26,619,384],{"class":36},[26,621,369],{"class":36},[26,623,49],{"class":36},[26,625,374],{"class":36},[26,627,52],{"class":36},[26,629,630],{"class":36}," -w",[26,632,633],{"class":36}," -i",[26,635,404],{"class":310},[26,637,638],{"class":36},"{\"file_url\": \"https:\u002F\u002Fexample.com\u002Freport.pdf\"}",[26,640,641],{"class":310},"'\n",[11,643,644,645,648,649,652],{},"Two fields in that response are worth reading before you design anything around it. The Markdown comes back as a link rather than as a string, and the link carries ",[23,646,647],{},"link_expires_at"," about an hour out with ",[23,650,651],{},"file_expires_at"," seven days out. Fetch the bytes in the same job that produced them. A pipeline that stores the URL and reads it next week has stored nothing.",[96,654,656],{"id":655},"the-scanned-case-and-how-you-know","The scanned case, and how you know",[11,658,659,660,565,663,666],{},"The same run returned ",[23,661,662],{},"ocr_ran: false",[23,664,665],{},"ocr_units: 0",", because that document had a text layer and nothing needed optical recognition.",[11,668,669,670,672,673,675],{},"That pair of fields is the honest answer to the scanned-PDF problem. You pass ",[23,671,474],{}," on documents that might be scans, and the response tells you afterwards whether it actually ran. The billing follows the same logic: the endpoint is TIERED, the base call is one credit, an OCR-executed call totals five, and passing ",[23,674,474],{}," holds the larger amount but settles the smaller one when the file turned out to have text after all.",[11,677,678],{},"So the safe default on a mixed folder is to turn OCR on and let the meter decide, rather than trying to detect scans yourself first. Detecting them costs a read anyway, and getting it wrong costs a silent empty document.",[96,680,682],{"id":681},"python-and-why-the-library-is-not-the-hard-part","Python, and why the library is not the hard part",[11,684,685],{},"The most common version of this question asks how to do it in Python, and the honest answer is that Python has good libraries for it. PyMuPDF, pdfplumber and marker all convert well, and marker in particular produces very clean Markdown.",[11,687,688,689,693],{},"What none of them ships is the OCR engine for the scanned pages, the layout model for multi-column academic papers, or somebody to keep both current. That is the same split we drew for web pages in ",[82,690,692],{"href":691},"\u002Fblog\u002Fguides\u002Fpython-web-scraping-without-a-scraper","Web Scraping in Python Without Maintaining a Scraper",": keep the code that is genuinely yours, and buy the part that is infrastructure. Calling the endpoint from Python is four lines and leaves your existing pipeline intact.",[96,695,697],{"id":696},"what-about-converting-for-claude-or-another-model","What about converting for Claude or another model?",[11,699,700,701,703],{},"Nothing special is required, which is the point of converting at all. Markdown is what every current model reads most reliably, so the same output serves a chunker, a retrieval index and a prompt you paste by hand. If the document is going straight into a context window rather than into an index, set ",[23,702,434],{}," to drop running headers and page furniture, and skip the chunking half of this guide entirely.",[88,705,707],{"id":706},"does-converting-to-markdown-actually-save-tokens","Does converting to Markdown actually save tokens?",[11,709,710],{},"Yes, and the size of the saving is the part people get wrong in both directions. Two separate builders on r\u002FLocalLLaMA and r\u002FLLMDevs shipped HTML-to-Markdown converters this year with the same headline claim, roughly two thirds fewer tokens than raw HTML. That number is believable for a content-heavy web page and misleading as a general rule.",[11,712,713],{},"The saving comes from deleting markup, not from compressing prose. So it is large for HTML, where tags and attributes can outweigh the text, and near zero for a DOCX whose content was already mostly words. If your corpus is web pages, the conversion pays for itself in embedding costs alone. If it is Office documents, convert for consistency rather than for tokens, and do not budget a saving that will not arrive.",[11,715,716],{},"The second-order effect is bigger than the token count anyway. Markdown keeps headings, lists and table structure as text, which means a chunker can split on semantic boundaries rather than on character counts. A chunk that starts at a heading retrieves better than a chunk that starts mid-sentence, and that improvement does not show up in a token comparison at all.",[88,718,720],{"id":719},"which-endpoint-should-i-use-for-which-job","Which endpoint should I use for which job?",[133,722,723,744],{},[136,724,725],{},[139,726,727,730,733,736,739,741],{},[142,728,729],{},"Endpoint",[142,731,732],{},"What it does",[142,734,735],{},"Input",[142,737,738],{},"Output",[142,740,203],{},[142,742,743],{},"Billing",[152,745,746,771,795,818,841,864,886],{},[139,747,748,754,757,760,765,768],{},[157,749,750],{},[82,751,752],{"href":345},[23,753,348],{},[157,755,756],{},"Convert a file to Markdown, 60+ formats",[157,758,759],{},"File URL up to 25MB",[157,761,762,763],{},"Markdown, plus ",[23,764,527],{},[157,766,767],{},"The main conversion step",[157,769,770],{},"Per call, higher only when OCR runs",[139,772,773,780,783,786,789,792],{},[157,774,775],{},[82,776,777],{"href":345},[23,778,779],{},"context.dev\u002Fweb\u002Fscrape\u002Fmarkdown",[157,781,782],{},"Convert a live web page to Markdown",[157,784,785],{},"URL",[157,787,788],{},"Markdown",[157,790,791],{},"Pages, not files",[157,793,794],{},"Per call",[139,796,797,804,807,810,813,816],{},[157,798,799],{},[82,800,801],{"href":345},[23,802,803],{},"context.dev\u002Fweb\u002Fcrawl",[157,805,806],{},"Follow links across a site",[157,808,809],{},"Start URL",[157,811,812],{},"One Markdown document per page",[157,814,815],{},"Ingesting a whole site",[157,817,794],{},[139,819,820,827,830,833,836,839],{},[157,821,822],{},[82,823,824],{"href":345},[23,825,826],{},"context.dev\u002Fweb\u002Fscrape\u002Fsitemap",[157,828,829],{},"Enumerate a site's URLs",[157,831,832],{},"Domain",[157,834,835],{},"URL list",[157,837,838],{},"Planning a crawl before running it",[157,840,794],{},[139,842,843,850,853,855,858,861],{},[157,844,845],{},[82,846,847],{"href":345},[23,848,849],{},"octen\u002Fextract",[157,851,852],{},"Clean Markdown from up to 20 URLs per call",[157,854,835],{},[157,856,857],{},"LLM-ready Markdown",[157,859,860],{},"Batches of known pages",[157,862,863],{},"Per result",[139,865,866,873,876,878,881,884],{},[157,867,868],{},[82,869,870],{"href":345},[23,871,872],{},"tinyfish\u002Ffetch",[157,874,875],{},"Full page text for up to 10 URLs",[157,877,835],{},[157,879,880],{},"Page text",[157,882,883],{},"Cheap bulk fetching",[157,885,794],{},[139,887,888,894,897,900,903,906],{},[157,889,890],{},[82,891,892],{"href":345},[23,893,480],{},[157,895,896],{},"OCR a loose image",[157,898,899],{},"Image",[157,901,902],{},"Text with a confidence score",[157,904,905],{},"Screenshots and photos",[157,907,794],{},[11,909,910,911,914],{},"Every row verified with ",[23,912,913],{},"monid inspect"," on 2026-08-20. The billing column is the shape rather than a figure, because per call and per result change how you batch and a price does not stay true.",[88,916,918],{"id":917},"what-does-an-ingestion-run-actually-cost","What does an ingestion run actually cost?",[11,920,921],{},"Less than the embeddings, in almost every case, which is why the conversion step deserves more attention than its budget line suggests.",[11,923,924],{},"A corpus of a few thousand mixed documents converts for single-digit dollars, because the main conversion endpoint bills per call rather than per page and most documents are one call. The number that moves is OCR, and only for the files that genuinely need it, since the higher rate settles only when OCR actually ran.",[11,926,927],{},"The comparison worth making is against the embedding bill rather than against doing it yourself. If a naive HTML extraction inflates a page from 689 characters to 74,552, you pay that inflation once in conversion and then again on every embedding and every retrieval that includes the diluted chunk. Cleaning at the door is the cheapest place in the pipeline to fix it.",[11,929,930,931,266],{},"Discovery and inspection are free, so the whole format list, every switch and the exact billing behaviour are readable before spending anything. That is the property that makes a metered balance suit an ingestion job: access costs nothing until it is used, so a one-off corpus load does not need a plan. Prices at ",[82,932,452],{"href":451},[934,935,938],"tool-cta",{"category":936,"title":937},"web","Convert the corpus before you chunk it",[11,939,940],{},"Discover what handles which format, read the switches, and see current prices. Nothing bills until you run.",[88,942,944],{"id":943},"when-should-you-not-use-monid","When should you not use Monid?",[11,946,947],{},"If the documents cannot leave your network, run it locally and accept the maintenance. Regulated corpora, client-confidential files and anything under a data residency commitment belong in a local pipeline, and the honest answer is that six format libraries and a native OCR build are the price of that constraint. No hosted endpoint solves a rule that says the bytes stay put.",[11,949,950],{},"If you have one format and one shape, use the library. A pipeline that only ever sees clean text-layer PDFs from one generator does not need a general conversion service. A well-chosen local parser will be faster and free.",[11,952,953],{},"If your volume is enormous and steady, the arithmetic changes. Per-call conversion is excellent for a corpus load and for a steady trickle of new documents, and it stops being the cheapest option somewhere above a sustained high rate where a self-hosted converter on your own compute wins on unit cost.",[11,955,956],{},"And if you need conversion accuracy guarantees for legal or medical documents, no general endpoint provides them. Specialist vendors sell validated extraction with an accuracy commitment attached, and that commitment, not the conversion, is what you would be buying.",[88,958,960],{"id":959},"conclusion","Conclusion",[11,962,963],{},"The best practice for ingesting mixed document types is to treat conversion as its own step with its own tests, rather than as a preamble to chunking. Convert everything through one door, turn OCR on across the whole batch because it only bills when it runs, assert on character counts and structure before embedding anything, and chunk last.",[11,965,966],{},"What matters more than the tool choice: the failures in this pipeline do not raise exceptions. A scanned PDF returns an empty string, a two-column paper returns interleaved text, and an HTML page returns a navigation menu. All three embed successfully and retrieve badly, and none of them appear in a log. Three assertions between conversion and chunking catch all three, and they are the cheapest code in the whole system.",[11,968,969,970,973,974,977,978,266],{},"The free next step costs nothing. Run ",[23,971,972],{},"monid discover -q \"parse pdf document to text\""," to see what exists, ",[23,975,976],{},"monid inspect -p context.dev -e \u002Fparse"," to read the full format list and the OCR billing behaviour, then one paid run on the ugliest ten documents in your corpus rather than the cleanest. Start at ",[82,979,982],{"href":980,"rel":981},"https:\u002F\u002Fmonid.ai",[246,247],"monid.ai",[88,984,986],{"id":985},"faq","FAQ",[988,989,991],"faq-item",{"q":990},"How do I handle scanned PDFs that have no text layer?",[11,992,993,994,998,999,1001],{},"Route them to OCR, and detect them by character count rather than by file inspection. A scanned page returns an empty or near-empty string from a text extractor, so a minimum-characters-per-page threshold identifies them reliably without you having to classify the batch in advance. With ",[82,995,996],{"href":345},[23,997,348],{}," you can set ",[23,1000,474],{}," across a mixed batch, because the OCR rate settles only on the files where OCR actually ran.",[988,1003,1005],{"q":1004},"How should I de-duplicate before embedding?",[11,1006,1007],{},"Hash the normalised text and compare strings, before anything reaches the embedding model. Vector-space near-duplicate detection is a real technique but it is the expensive way to catch the common case, which is the same document appearing twice under two filenames. Normalise whitespace, strip the Markdown, hash, and drop exact matches first, then use similarity only for the residue.",[988,1009,1011],{"q":1010},"Where should chunking happen, before or after conversion?",[11,1012,1013],{},"After, always. Chunking operates on text and conversion produces the text, so chunking first means chunking whatever raw bytes you had. The more useful version of the question is what to chunk on, and converting to Markdown first is what makes the good answer available: split on headings and list boundaries rather than on a character count, which is only possible once the structure survives as text.",[988,1015,1017],{"q":1016},"Can I keep documents out of a vendor's logs?",[11,1018,1019,1020,1023,1024,1026],{},"Sometimes, and it is a field rather than a conversation. The parse endpoint exposes a ",[23,1021,1022],{},"zdr"," switch that bypasses the vendor's shared caches and omits request and response content from its retained usage logs, though it requires zero data retention to be enabled on the vendor account first and fails explicitly if it is not. Read that field's behaviour with ",[23,1025,913],{}," before assuming it applies to your account.",[11,1028,1029],{},[1030,1031,1032],"em",{},"Last updated August 2026.",[1034,1035,1036],"style",{},"html pre.shiki code .s2Zo4, html code.shiki .s2Zo4{--shiki-light:#6182B8;--shiki-default:#82AAFF;--shiki-dark:#82AAFF}html pre.shiki code .sfazB, html code.shiki .sfazB{--shiki-light:#91B859;--shiki-default:#C3E88D;--shiki-dark:#C3E88D}html .light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html.light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .sBMFI, html code.shiki .sBMFI{--shiki-light:#E2931D;--shiki-default:#FFCB6B;--shiki-dark:#FFCB6B}html pre.shiki code .sMK4o, html code.shiki .sMK4o{--shiki-light:#39ADB5;--shiki-default:#89DDFF;--shiki-dark:#89DDFF}html pre.shiki code .sTEyZ, html code.shiki .sTEyZ{--shiki-light:#90A4AE;--shiki-default:#EEFFFF;--shiki-dark:#BABED8}",{"title":21,"searchDepth":295,"depth":295,"links":1038},[1039,1045,1052,1058,1059,1060,1061,1062,1063],{"id":90,"depth":295,"text":91,"children":1040},[1041,1042,1043,1044],{"id":98,"depth":398,"text":99},{"id":108,"depth":398,"text":109},{"id":118,"depth":398,"text":119},{"id":130,"depth":398,"text":131},{"id":230,"depth":295,"text":231,"children":1046},[1047,1048,1049,1050,1051],{"id":237,"depth":398,"text":238},{"id":269,"depth":398,"text":270},{"id":330,"depth":398,"text":331},{"id":455,"depth":398,"text":456},{"id":539,"depth":398,"text":540},{"id":572,"depth":295,"text":573,"children":1053},[1054,1055,1056,1057],{"id":579,"depth":398,"text":580},{"id":655,"depth":398,"text":656},{"id":681,"depth":398,"text":682},{"id":696,"depth":398,"text":697},{"id":706,"depth":295,"text":707},{"id":719,"depth":295,"text":720},{"id":917,"depth":295,"text":918},{"id":943,"depth":295,"text":944},{"id":959,"depth":295,"text":960},{"id":985,"depth":295,"text":986},"Search & RAG","\u002Fimg\u002Fblog\u002Fingest-mixed-documents-llm-embedding.png","PDFs, Office files and HTML all become one clean format before chunking. Where the pipeline actually breaks, and which endpoint handles which format.",false,"md",null,"\u002Fimg\u002Fblog\u002Fingest-mixed-documents-llm-embedding-card.png",{},true,"\u002Fblog\u002Fguides\u002Fingest-mixed-documents-llm-embedding","2026-08-20","11 min",{"title":5,"description":1066},"blog\u002Fguides\u002Fingest-mixed-documents-llm-embedding",[1079,1080,1081,1082,1083],"rag","embeddings","pdf","ocr","markdown","2026-08-26","iwUtenJ7EyHP02YPZqd_S-VMe9oOD_8pXQliJYG9rqk",[1087,1102,1115,1129,1141,1154,1166,1180,1193,1205,1218,1230,1242,1254,1267,1279,1291,1304,1317,1329,1342,1354,1366,1379,1391,1403,1415,1427,1439,1452,1463,1475,1488,1501,1512,1524,1537,1547,1560,1571,1582,1595,1607,1618,1630,1640,1653,1665,1675,1687,1699,1711,1722,1734,1746,1758,1771,1784,1796,1807,1819,1831,1844,1856,1869,1881,1892,1903,1916,1926,1937,1947,1959,1969,1980,1991,2004,2014,2025,2033,2043,2053,2064,2076,2088,2098,2111,2121,2133,2143,2152,2162,2172,2180,2190,2199,2210,2220,2222,2233,2245,2252,2262,2270,2280,2289,2299,2310,2317,2325,2334,2345,2356,2367,2375,2385,2396,2405,2414,2424,2434,2442,2455,2465,2475,2486,2496,2506,2516,2525,2535,2542,2552,2564,2573,2586,2595,2604,2616,2625,2634,2642,2651,2660],{"path":1088,"title":1089,"description":1090,"publishedAt":1091,"author":6,"category":1092,"tags":1093,"readTime":1099,"cover":1100,"listingCover":1101,"image":1069},"\u002Fblog\u002Fguides\u002Fbacklink-api-what-it-returns","Backlink API: 74% of the Links Are Already Gone","One domain reported 284 million backlinks all time and 74 million live. The anchor list opened with black-hat spam. Which of the three calls to make.","2026-10-03","Product",[1094,1095,1096,1097,1098],"backlink api","referring domains api","anchor text api","link data","backlink checker","10 min","\u002Fimg\u002Fblog\u002Fbacklink-api-what-it-returns.png","\u002Fimg\u002Fblog\u002Fbacklink-api-what-it-returns-card.png",{"path":1103,"title":1104,"description":1105,"publishedAt":1091,"author":6,"category":1106,"tags":1107,"readTime":1099,"cover":1113,"listingCover":1114,"image":1069},"\u002Fblog\u002Fguides\u002Fglassdoor-api-employee-reviews","Glassdoor API: Asking for 10 Reviews Bills You for 50","Same endpoint, two calls. limit 10 returned 10 reviews and billed 50 units. limit 100 returned 100 and billed 100. And eight ratings use two scales.","Sales & enrichment",[1108,1109,1110,1111,1112],"glassdoor api","employee reviews api","employer brand data","company reviews","per unit billing","\u002Fimg\u002Fblog\u002Fglassdoor-api-employee-reviews.png","\u002Fimg\u002Fblog\u002Fglassdoor-api-employee-reviews-card.png",{"path":1116,"title":1117,"description":1118,"publishedAt":1119,"author":6,"category":1120,"tags":1121,"readTime":1099,"cover":1127,"listingCover":1128,"image":1069},"\u002Fblog\u002Fguides\u002Faddress-validation-api","Address Validation API: One Said Invalid, One Said High","Same address, same minute, two endpoints. One returned valid false and confidence 0 while handing back a perfect parse. The other returned high.","2026-10-02","Local data",[1122,1123,1124,1125,1126],"address validation api","address parser api","address standardization","geocoding api","postal address","\u002Fimg\u002Fblog\u002Faddress-validation-api.png","\u002Fimg\u002Fblog\u002Faddress-validation-api-card.png",{"path":1130,"title":1131,"description":1132,"publishedAt":1119,"author":6,"category":1120,"tags":1133,"readTime":1099,"cover":1139,"listingCover":1140,"image":1069},"\u002Fblog\u002Fguides\u002Ftripadvisor-api-search","TripAdvisor API: 2,000 Results, 30 a Page, One Id","A search for Kyoto hotels returned 30 typed rows out of 2,000, every one carrying a rating and a place id. What the typed fields are actually for.",[1134,1135,1136,1137,1138],"tripadvisor api","tripadvisor scraper","hotel data api","travel data","attraction data","\u002Fimg\u002Fblog\u002Ftripadvisor-api-search.png","\u002Fimg\u002Fblog\u002Ftripadvisor-api-search-card.png",{"path":1142,"title":1143,"description":1144,"publishedAt":1145,"author":6,"category":1106,"tags":1146,"readTime":1099,"cover":1152,"listingCover":1153,"image":1069},"\u002Fblog\u002Fguides\u002Findeed-api-job-search","Indeed API: Passing false to One Filter Returned a 502, Twice","Omit the flag and you get 245 jobs with 38 duplicates. Switch it on and you get 234 with 19. Send it as false and the call fails without charging.","2026-10-01",[1147,1148,1149,1150,1151],"indeed api","indeed job search api","job search api","hiring data","duplicate job postings","\u002Fimg\u002Fblog\u002Findeed-api-job-search.png","\u002Fimg\u002Fblog\u002Findeed-api-job-search-card.png",{"path":1155,"title":1156,"description":1157,"publishedAt":1145,"author":6,"category":1120,"tags":1158,"readTime":1099,"cover":1164,"listingCover":1165,"image":1069},"\u002Fblog\u002Fguides\u002Fyelp-api-business-search","Yelp API: The Four Ads Above the Results Had 5, 75, 80 and 199 Reviews","The top organic result had 2,156. The ads sitting above it had a fraction of that. Both arrive in the same response, in two separate arrays.",[1159,1160,1161,1162,1163],"yelp api","yelp scraper","local business data","yelp reviews api","local lead list","\u002Fimg\u002Fblog\u002Fyelp-api-business-search.png","\u002Fimg\u002Fblog\u002Fyelp-api-business-search-card.png",{"path":1167,"title":1168,"description":1169,"publishedAt":1170,"author":6,"category":1064,"tags":1171,"readTime":1177,"cover":1178,"listingCover":1179,"image":1069},"\u002Fblog\u002Fguides\u002Fanswer-engine-optimization-four-surfaces","Answer Engine Optimization: We Asked Four Surfaces the Same Six Questions","Four AI answer surfaces, six questions, one day, 24 answers. Only two questions had a vendor all four agreed on, and we were named in zero of them.","2026-09-28",[1172,1173,1174,1175,1176],"answer engine optimization","geo seo","llm seo","generative engine optimization","ai answers","12 min","\u002Fimg\u002Fblog\u002Fanswer-engine-optimization-four-surfaces.png","\u002Fimg\u002Fblog\u002Fanswer-engine-optimization-four-surfaces-card.png",{"path":1181,"title":1182,"description":1183,"publishedAt":1170,"author":6,"category":1184,"tags":1185,"readTime":1099,"cover":1191,"listingCover":1192,"image":1069},"\u002Fblog\u002Fguides\u002Fdownload-youtube-thumbnail","Download a YouTube Thumbnail: Which of the Six URLs Actually Exists","Six predictable thumbnail URLs, four videos, checked one at a time. Three names answered every time and three returned 404 on the oldest upload.","Social data",[1186,1187,1188,1189,1190],"download youtube thumbnail","youtube thumbnail url","maxresdefault","youtube thumbnail api","video metadata","\u002Fimg\u002Fblog\u002Fdownload-youtube-thumbnail.png","\u002Fimg\u002Fblog\u002Fdownload-youtube-thumbnail-card.png",{"path":1194,"title":1195,"description":1196,"publishedAt":1170,"author":6,"category":1064,"tags":1197,"readTime":1075,"cover":1203,"listingCover":1204,"image":1069},"\u002Fblog\u002Fguides\u002Fperplexity-api-two-surfaces","Perplexity API: Sonar and the App Answered the Same Question Differently","Same six questions, same day. Sonar replied in about three seconds with 84 cited domains. The consumer app took up to 92 and named different vendors.",[1198,1199,1200,1201,1202],"perplexity api","perplexity sonar api","sonar model","ai answer api","perplexity citations","\u002Fimg\u002Fblog\u002Fperplexity-api-two-surfaces.png","\u002Fimg\u002Fblog\u002Fperplexity-api-two-surfaces-card.png",{"path":1206,"title":1207,"description":1208,"publishedAt":1209,"author":6,"category":1092,"tags":1210,"readTime":1075,"cover":1216,"listingCover":1217,"image":1069},"\u002Fblog\u002Fguides\u002Fllm-as-a-judge-typesafe-jev","LLM as a Judge, or a Model That Only Judges: Jev Measured","Four typed judgments on one review in 1.9 seconds, billed on 631 input tokens with output free. The severity score came back 1.1, and it is zero-indexed.","2026-09-24",[1211,1212,1213,1214,1215],"llm as a judge","typesafe jev","classification api","structured judgment","system one","\u002Fimg\u002Fblog\u002Fllm-as-a-judge-typesafe-jev.png","\u002Fimg\u002Fblog\u002Fllm-as-a-judge-typesafe-jev-card.png",{"path":1219,"title":1220,"description":1221,"publishedAt":1209,"author":6,"category":1092,"tags":1222,"readTime":1099,"cover":1228,"listingCover":1229,"image":1069},"\u002Fblog\u002Fguides\u002Fmeta-muse-agent-tools","Meta Muse Builds Its Own Tools: Here It Does Not Have To","Muse writes a connector for anything with an API. One spec of 38 paths and 42 operations means it writes one connector instead of two thousand.",[1223,1224,1225,1226,1227],"meta muse","muse agent tools","agent connectors","openapi for agents","rest api for ai agents","\u002Fimg\u002Fblog\u002Fmeta-muse-agent-tools.png","\u002Fimg\u002Fblog\u002Fmeta-muse-agent-tools-card.png",{"path":1231,"title":1232,"description":1233,"publishedAt":1209,"author":6,"category":1106,"tags":1234,"readTime":1075,"cover":1240,"listingCover":1241,"image":1069},"\u002Fblog\u002Fguides\u002Fsales-automation-by-the-call","Sales Automation Without a Seat: Four Calls and One Receipt","The automation most teams buy is four lookups in a row. Here is each one measured, including the one that charged us for a record with a null title.",[1235,1236,1237,1238,1239],"sales automation","what is sales automation","sales workflow","outbound automation","lead generation","\u002Fimg\u002Fblog\u002Fsales-automation-by-the-call.png","\u002Fimg\u002Fblog\u002Fsales-automation-by-the-call-card.png",{"path":1243,"title":1244,"description":1245,"publishedAt":1209,"author":6,"category":1092,"tags":1246,"readTime":1075,"cover":1252,"listingCover":1253,"image":1069},"\u002Fblog\u002Fguides\u002Fzapier-alternatives-tool-layer","Zapier Alternatives: The Directory's Own List Had Trello On It","We pulled a directory list of Zapier alternatives. The array named alternatives held Trello, Jira and Notion. The competitors were in another field.",[1247,1248,1249,1250,1251],"zapier alternatives","n8n alternatives","n8n api","workflow automation","software alternatives api","\u002Fimg\u002Fblog\u002Fzapier-alternatives-tool-layer.png","\u002Fimg\u002Fblog\u002Fzapier-alternatives-tool-layer-card.png",{"path":1255,"title":1256,"description":1257,"publishedAt":1258,"author":6,"category":1106,"tags":1259,"readTime":1177,"cover":1265,"listingCover":1266,"image":1069},"\u002Fblog\u002Fguides\u002Fai-lead-score-calculation-method","AI Lead Score Calculation: The Inputs Disagree Before the Model Runs","Two providers gave one company 580 and 1,011 employees. A third resolved a famous domain to a different company with a plausible record.","2026-09-23",[1260,1261,1262,1263,1264],"ai lead score calculation method","lead scoring","icp scoring","lead qualification","firmographics","\u002Fimg\u002Fblog\u002Fai-lead-score-calculation-method.png","\u002Fimg\u002Fblog\u002Fai-lead-score-calculation-method-card.png",{"path":1268,"title":1269,"description":1270,"publishedAt":1258,"author":6,"category":1092,"tags":1271,"readTime":1075,"cover":1277,"listingCover":1278,"image":1069},"\u002Fblog\u002Fguides\u002Fkeyword-cannibalization-check","Keyword Cannibalization: We Found Six of Our Own Pages Fighting","A site-restricted search returned six of our guides for one concept. The unrestricted search returned none of them. That gap is the problem.",[1272,1273,1274,1275,1276],"keyword cannibalization","ahrefs keyword cannibalization","content audit","serp api","internal competition","\u002Fimg\u002Fblog\u002Fkeyword-cannibalization-check.png","\u002Fimg\u002Fblog\u002Fkeyword-cannibalization-check-card.png",{"path":1280,"title":1281,"description":1282,"publishedAt":1258,"author":6,"category":1106,"tags":1283,"readTime":1075,"cover":1289,"listingCover":1290,"image":1069},"\u002Fblog\u002Fguides\u002Ftarget-company-url-analysis","Target Company URL Analysis: What One Domain Tells You in Four Calls","Start with nothing but a URL. Four calls returned the company, its size, its stack with checkable evidence, and 445 known contacts split by department.",[1284,1285,1286,1287,1288],"target company url analysis","company url lookup","domain to company","website analysis api","account research","\u002Fimg\u002Fblog\u002Ftarget-company-url-analysis.png","\u002Fimg\u002Fblog\u002Ftarget-company-url-analysis-card.png",{"path":1292,"title":1293,"description":1294,"publishedAt":1295,"author":6,"category":1106,"tags":1296,"readTime":1075,"cover":1302,"listingCover":1303,"image":1069},"\u002Fblog\u002Fguides\u002Fapollo-extension-or-api","Apollo Extension or Apollo API: Which One Does the Job You Have","One company call returned 66 fields including 271 technologies in 2.5 seconds. A person call returned a null title and a null email, and still billed.","2026-09-21",[1297,1298,1299,1300,1301],"apollo.io extension","apollo extension","apollo chrome extension","apollo api","apollo lead generation","\u002Fimg\u002Fblog\u002Fapollo-extension-or-api.png","\u002Fimg\u002Fblog\u002Fapollo-extension-or-api-card.png",{"path":1305,"title":1306,"description":1307,"publishedAt":1295,"author":6,"category":1308,"tags":1309,"readTime":1177,"cover":1315,"listingCover":1316,"image":1069},"\u002Fblog\u002Fguides\u002Fsearch-for-upc-barcode-lookup-api","UPC Lookup by API: Two Databases, and One Barcode Neither Had","A soda barcode returned ingredients and a nutrition grade. A book came from another database. A phone cable came back not found, and still billed.","Ecommerce",[1310,1311,1312,1313,1314],"upc lookup","upc code lookup","barcode lookup api","search for upc","gtin lookup","\u002Fimg\u002Fblog\u002Fsearch-for-upc-barcode-lookup-api.png","\u002Fimg\u002Fblog\u002Fsearch-for-upc-barcode-lookup-api-card.png",{"path":1318,"title":1319,"description":1320,"publishedAt":1295,"author":6,"category":1184,"tags":1321,"readTime":1075,"cover":1327,"listingCover":1328,"image":1069},"\u002Fblog\u002Fguides\u002Fsocial-searcher-username-across-platforms","Social Searcher by API: One Username, Six Platforms, Six Schemas","One handle through one endpoint on six platforms. The identity field changed name three times and one platform returned a profile that was not the brand.",[1322,1323,1324,1325,1326],"social searcher","username search","instant username search","social profile","cross platform social search","\u002Fimg\u002Fblog\u002Fsocial-searcher-username-across-platforms.png","\u002Fimg\u002Fblog\u002Fsocial-searcher-username-across-platforms-card.png",{"path":1330,"title":1331,"description":1332,"publishedAt":1333,"author":6,"category":1092,"tags":1334,"readTime":1075,"cover":1340,"listingCover":1341,"image":1069},"\u002Fblog\u002Fguides\u002Fahrefs-vs-kwfinder-link-half","Ahrefs vs KWFinder: Both Halves Are Available by the Call","We said the keyword half was not buyable per call. That was wrong. Here is the corrected answer, and the 740x volume disagreement we found while checking.","2026-09-18",[1335,1336,1337,1338,1339],"ahrefs vs kwfinder","kwfinder vs ahrefs","keyword research api","ahrefs alternative","seo api by the call","\u002Fimg\u002Fblog\u002Fahrefs-vs-kwfinder-link-half.png","\u002Fimg\u002Fblog\u002Fahrefs-vs-kwfinder-link-half-card.png",{"path":1343,"title":1344,"description":1345,"publishedAt":1333,"author":6,"category":1184,"tags":1346,"readTime":1075,"cover":1352,"listingCover":1353,"image":1069},"\u002Fblog\u002Fguides\u002Ffacebook-profile-viewer-api","Facebook Profile Viewer: What an API Returns, and What It Will Not","The scraper whose docs say it refuses personal profiles returned one anyway. Its two follower counts for that person differed by one.",[1347,1348,1349,1350,1351],"facebook profile viewer","facebook profile api","facebook page scraper","public facebook data","facebook api","\u002Fimg\u002Fblog\u002Ffacebook-profile-viewer-api.png","\u002Fimg\u002Fblog\u002Ffacebook-profile-viewer-api-card.png",{"path":1355,"title":1356,"description":1357,"publishedAt":1333,"author":6,"category":1106,"tags":1358,"readTime":1075,"cover":1364,"listingCover":1365,"image":1069},"\u002Fblog\u002Fguides\u002Fh1b-data-api-sponsorship-records","H-1B Data by API: 549,133 Records, and the Domain Field Is Invented","One call returned half a million sponsorship records. The top wage was a React developer on 1.9 million. The company domain on every row does not resolve.",[1359,1360,1361,1362,1363],"h1b data","h1b api","h1b salary data","visa sponsorship data","labor condition application","\u002Fimg\u002Fblog\u002Fh1b-data-api-sponsorship-records.png","\u002Fimg\u002Fblog\u002Fh1b-data-api-sponsorship-records-card.png",{"path":1367,"title":1368,"description":1369,"publishedAt":1370,"author":6,"category":1064,"tags":1371,"readTime":1075,"cover":1377,"listingCover":1378,"image":1069},"\u002Fblog\u002Fguides\u002Farxiv-api-paper-search","arXiv API for Paper Search: Three of Twenty Results Were Off Topic","Restricting a paper search to arXiv returned three papers about cognitive augmentation instead of retrieval augmentation. The blended source returned none.","2026-09-17",[1372,1373,1374,1375,1376],"arxiv api","paper search api","academic search api","literature review api","semantic scholar","\u002Fimg\u002Fblog\u002Farxiv-api-paper-search.png","\u002Fimg\u002Fblog\u002Farxiv-api-paper-search-card.png",{"path":1380,"title":1381,"description":1382,"publishedAt":1370,"author":6,"category":1106,"tags":1383,"readTime":1075,"cover":1389,"listingCover":1390,"image":1069},"\u002Fblog\u002Fguides\u002Fbuzzfile-search-by-name-company-lookup","Buzzfile and the Company Lookup API: One Name, Ten Companies","What Buzzfile is, and what a name search returns by API. One free call on Patagonia gave ten companies. A rival fuzzy search gave two, both wrong.",[1384,1385,1386,1387,1388],"buzzfile","buzzfile search by name","company lookup api","name to domain","company search by name","\u002Fimg\u002Fblog\u002Fbuzzfile-search-by-name-company-lookup.png","\u002Fimg\u002Fblog\u002Fbuzzfile-search-by-name-company-lookup-card.png",{"path":1392,"title":1393,"description":1394,"publishedAt":1370,"author":6,"category":1106,"tags":1395,"readTime":1075,"cover":1401,"listingCover":1402,"image":1069},"\u002Fblog\u002Fguides\u002Frocketreach-pricing-per-contact-lookups","RocketReach Pricing or Per-Contact Lookups: Where the Charge Lands","We paid the base charge for one executive and got an id, a country and a null title. A second provider returned the right job title and zero contacts.",[1396,1397,1398,1399,1400],"rocketreach pricing","rocketreach","contact lookup api","people enrichment pricing","credits","\u002Fimg\u002Fblog\u002Frocketreach-pricing-per-contact-lookups.png","\u002Fimg\u002Fblog\u002Frocketreach-pricing-per-contact-lookups-card.png",{"path":1404,"title":1405,"description":1406,"publishedAt":1407,"author":6,"category":1106,"tags":1408,"readTime":1075,"cover":1413,"listingCover":1414,"image":1069},"\u002Fblog\u002Fguides\u002Fb2b-prospecting-tool-api","B2B Prospecting Tool by API: The List Is Free, the Name Is the Bill","One free call returned 100 sales leaders at 47 companies with names redacted. Revealing who they are is the paid step. How three prospecting routes bill.","2026-09-16",[1409,1410,1411,1412,1239],"b2b prospecting tool","b2b leads","b2b sms outreach","prospecting api","\u002Fimg\u002Fblog\u002Fb2b-prospecting-tool-api.png","\u002Fimg\u002Fblog\u002Fb2b-prospecting-tool-api-card.png",{"path":1416,"title":1417,"description":1418,"publishedAt":1407,"author":6,"category":1106,"tags":1419,"readTime":1075,"cover":1425,"listingCover":1426,"image":1069},"\u002Fblog\u002Fguides\u002Fcrunchbase-pro-vs-api-by-the-call","Crunchbase Pro or Crunchbase by the Call: What Each Actually Returns","Sixty-nine organisations match Stripe. The profile carried seven fields and no funding total. Twenty-five rounds came back with types and no amounts.",[1420,1421,1422,1423,1424],"crunchbase pro","crunchbase api","crunchbase pricing","funding rounds api","company data","\u002Fimg\u002Fblog\u002Fcrunchbase-pro-vs-api-by-the-call.png","\u002Fimg\u002Fblog\u002Fcrunchbase-pro-vs-api-by-the-call-card.png",{"path":1428,"title":1429,"description":1430,"publishedAt":1407,"author":6,"category":1064,"tags":1431,"readTime":1075,"cover":1437,"listingCover":1438,"image":1069},"\u002Fblog\u002Fguides\u002Ftavily-api-key-or-one-key","Tavily API, or One Key: Four Search Engines, One Query, Zero URLs in Common","What the Tavily API is, and what four search engines did with one query in the same minute. Two shared no URLs at all. One found our own page.",[1432,1433,1434,1435,1436],"tavily api","tavily","tavily api key","web search api for agents","search api comparison","\u002Fimg\u002Fblog\u002Ftavily-api-key-or-one-key.png","\u002Fimg\u002Fblog\u002Ftavily-api-key-or-one-key-card.png",{"path":1440,"title":1441,"description":1442,"publishedAt":1443,"author":6,"category":1064,"tags":1444,"readTime":1075,"cover":1450,"listingCover":1451,"image":1069},"\u002Fblog\u002Fguides\u002Fahrefs-api-domain-rating-backlinks","Ahrefs API by the Call: Domain Rating, Backlinks, Paid Pages","We pulled our own backlink profile through the Ahrefs API: ten referring domains, all DR 91 or higher, six dofollow links out of thirty-three.","2026-09-15",[1445,1446,1447,1448,1449],"ahrefs api","ahrefs api documentation","domain rating api","backlinks api","seo data api","\u002Fimg\u002Fblog\u002Fahrefs-api-domain-rating-backlinks.png","\u002Fimg\u002Fblog\u002Fahrefs-api-domain-rating-backlinks-card.png",{"path":1453,"title":1454,"description":1455,"publishedAt":1443,"author":6,"category":1106,"tags":1456,"readTime":1075,"cover":1461,"listingCover":1462,"image":1069},"\u002Fblog\u002Fguides\u002Fapollo-email-finder-api","Apollo Email Finder API: Two Providers, One Mailbox, Two Owners","Apollo says patrick@stripe.com is the CEO, verified. Hunter.io says it is an IT administrator, observed. Same address, one day apart, both confident.",[1457,1300,1458,1459,1460],"apollo email finder","apollo.io api","people match","email to person","\u002Fimg\u002Fblog\u002Fapollo-email-finder-api.png","\u002Fimg\u002Fblog\u002Fapollo-email-finder-api-card.png",{"path":1464,"title":1465,"description":1466,"publishedAt":1443,"author":6,"category":1092,"tags":1467,"readTime":1099,"cover":1473,"listingCover":1474,"image":1069},"\u002Fblog\u002Fguides\u002Fdomain-name-availability-api","Domain Age Checker: One TLD Returns Nothing and Still Bills","Eleven domains through a domain age endpoint. Three TLDs came back empty with HTTP 200 and were billed. Two were refused for free.",[1468,1469,1470,1471,1472],"domain age checker","domain availability api","whois api","domain lookup","domain fraud signal","\u002Fimg\u002Fblog\u002Fdomain-name-availability-api.png","\u002Fimg\u002Fblog\u002Fdomain-name-availability-api-card.png",{"path":1476,"title":1477,"description":1478,"publishedAt":1443,"author":6,"category":1064,"tags":1479,"readTime":1177,"cover":1486,"listingCover":1487,"image":1069},"\u002Fblog\u002Fguides\u002Ffirecrawl-with-claude-mcp-or-one-key","Claude MCP for Web Data: One Server Per Tool, or One Key?","One call mapped 3,779 documentation pages in eight seconds. The question is whether each tool that does this needs its own MCP server in your config.",[1480,1481,1482,1483,1484,1485],"claude mcp","claude mcp server","firecrawl mcp","firecrawl claude","mcp server","web scraping for agents","\u002Fimg\u002Fblog\u002Ffirecrawl-with-claude-mcp-or-one-key.png","\u002Fimg\u002Fblog\u002Ffirecrawl-with-claude-mcp-or-one-key-card.png",{"path":1489,"title":1490,"description":1491,"publishedAt":1443,"author":6,"category":1492,"tags":1493,"readTime":1075,"cover":1499,"listingCover":1500,"image":1069},"\u002Fblog\u002Fguides\u002Fhow-to-use-seedance-2-0-api","How to Use Seedance 2.0 by API: One Prompt, Two Models, Measured","Seedance 2.0 and 2.0 Mini rendered one 5-second prompt in the same two minutes for the same 108,900 tokens. One billed exactly twice the other.","Generative media",[1494,1495,1496,1497,1498],"how to use seedance 2.0","seedance 2.0","seedance api","text to video api","ai video generation","\u002Fimg\u002Fblog\u002Fhow-to-use-seedance-2-0-api.png","\u002Fimg\u002Fblog\u002Fhow-to-use-seedance-2-0-api-card.png",{"path":1502,"title":1503,"description":1504,"publishedAt":1443,"author":6,"category":1064,"tags":1505,"readTime":1075,"cover":1510,"listingCover":1511,"image":1069},"\u002Fblog\u002Fguides\u002Focten-ai-search-api","Octen AI Search API: 78 ms Inside, One Second End to End","Octen reports its own latency in every response. On our calls it said 78 ms; the round trip was a second. What search, extract and embedding return.",[1506,1507,1508,1509,1435],"octen ai","octen","octen search","octen api","\u002Fimg\u002Fblog\u002Focten-ai-search-api.png","\u002Fimg\u002Fblog\u002Focten-ai-search-api-card.png",{"path":1513,"title":1514,"description":1515,"publishedAt":1443,"author":6,"category":1184,"tags":1516,"readTime":1075,"cover":1522,"listingCover":1523,"image":1069},"\u002Fblog\u002Fguides\u002Freddit-user-information-extraction-api","Reddit User Information Extraction: Profile, Posts and Comments by API","One Reddit account, three endpoints: a profile with karma split two ways, a complete post history in one call, and comments that carry their thread.",[1517,1518,1519,1520,1521],"reddit user information extraction","reddit user api","reddit profile scraper","reddit comments api","social data","\u002Fimg\u002Fblog\u002Freddit-user-information-extraction-api.png","\u002Fimg\u002Fblog\u002Freddit-user-information-extraction-api-card.png",{"path":1525,"title":1526,"description":1527,"publishedAt":1528,"author":6,"category":1106,"tags":1529,"readTime":1075,"cover":1535,"listingCover":1536,"image":1069},"\u002Fblog\u002Fguides\u002Fhunter-io-api-domain-search-vs-email-finder","Hunter.io API: Domain Search Finds, Email Finder Guesses","Hunter.io domain search returned ten verified addresses at confidence 85. Email finder returned one address at confidence 16 with no source. Same provider.","2026-09-14",[1530,1531,1532,1533,1534],"hunter.io","hunter io api","email finder","domain search","email verification","\u002Fimg\u002Fblog\u002Fhunter-io-api-domain-search-vs-email-finder.png","\u002Fimg\u002Fblog\u002Fhunter-io-api-domain-search-vs-email-finder-card.png",{"path":1538,"title":1539,"description":1540,"publishedAt":1528,"author":6,"category":1106,"tags":1541,"readTime":1099,"cover":1545,"listingCover":1546,"image":1069},"\u002Fblog\u002Fguides\u002Freverse-email-lookup-api","Reverse Email Lookup API: What One Address Actually Returned","A reverse email lookup on a Stripe address returned an IT administrator, not the CEO. That is the endpoint working, and the fuzzy flag says why.",[1542,1543,1460,1544,1530],"reverse email lookup","reverse email lookup free","people enrichment","\u002Fimg\u002Fblog\u002Freverse-email-lookup-api.png","\u002Fimg\u002Fblog\u002Freverse-email-lookup-api-card.png",{"path":1548,"title":1549,"description":1550,"publishedAt":1528,"author":6,"category":1184,"tags":1551,"readTime":1177,"cover":1558,"listingCover":1559,"image":1069},"\u002Fblog\u002Fguides\u002Ftiktok-user-finder-api","TikTok Profile Viewer and Account Finder: What an API Returns","One profile call returned two different follower counts in one response. One keyword returned twenty accounts, ten of them regional handles of one brand.",[1552,1553,1554,1555,1556,1557],"tiktok profile viewer","tiktok account viewer","tiktok account finder","tiktok user search","tiktok api","tikhub","\u002Fimg\u002Fblog\u002Ftiktok-user-finder-api.png","\u002Fimg\u002Fblog\u002Ftiktok-user-finder-api-card.png",{"path":1561,"title":1562,"description":1563,"publishedAt":1564,"author":6,"category":1064,"tags":1565,"readTime":1075,"cover":1569,"listingCover":1570,"image":1069},"\u002Fblog\u002Fguides\u002Ffree-web-search-api-for-agents","TinyFish Made a Free Web Search API. One Query Is a Third.","We asked TinyFish Search one question five ways. It returned 27 pages, and the best single phrasing found 9 of them. Every call was free.","2026-09-13",[1566,1567,1079,1568],"web search","ai agents","retrieval","\u002Fimg\u002Fblog\u002Ffree-web-search-api-for-agents.png","\u002Fimg\u002Fblog\u002Ffree-web-search-api-for-agents-card.png",{"path":1572,"title":1573,"description":1574,"publishedAt":1575,"author":6,"category":1092,"tags":1576,"readTime":1075,"cover":1580,"listingCover":1581,"image":1069},"\u002Fblog\u002Fguides\u002Fai-agent-development-company-or-in-house","AI Agent Development Company or In House? The Hard Part","Two vendors, one domain, two different companies. One reported 2,930 employees and the other 63. Neither returned an error. That is the build.","2026-09-08",[1567,1577,1578,1579],"build vs buy","data quality","integration","\u002Fimg\u002Fblog\u002Fai-agent-development-company-or-in-house.png","\u002Fimg\u002Fblog\u002Fai-agent-development-company-or-in-house-card.png",{"path":1583,"title":1584,"description":1585,"publishedAt":1575,"author":6,"category":1586,"tags":1587,"readTime":1075,"cover":1593,"listingCover":1594,"image":1069},"\u002Fblog\u002Fguides\u002Ffind-companies-running-google-ads","How Do You Find Prospects Running Google Ads?","One landing page carried 1,537 distinct ad creatives. That count, not traffic, separates a company testing paid search from one committed to it.","Ads intelligence",[1588,1589,1590,1591,1592],"google ads","ads intelligence","prospecting","paid search","agency lead generation","\u002Fimg\u002Fblog\u002Ffind-companies-running-google-ads.png","\u002Fimg\u002Fblog\u002Ffind-companies-running-google-ads-card.png",{"path":1596,"title":1597,"description":1598,"publishedAt":1575,"author":6,"category":1308,"tags":1599,"readTime":1075,"cover":1605,"listingCover":1606,"image":1069},"\u002Fblog\u002Fguides\u002Fgoogle-shopping-api-lowest-price","Google Shopping API: The Lowest Price Is a Different Product","One search returned the same headphones from 22.99 to 399.99. The cheapest was a weekly rental. Sorting by price picks the row that is not the product.",[1600,1601,1602,1603,1604],"google shopping api","price comparison","ecommerce data","product matching","repricing","\u002Fimg\u002Fblog\u002Fgoogle-shopping-api-lowest-price.png","\u002Fimg\u002Fblog\u002Fgoogle-shopping-api-lowest-price-card.png",{"path":1608,"title":1609,"description":1610,"publishedAt":1575,"author":6,"category":1184,"tags":1611,"readTime":1075,"cover":1616,"listingCover":1617,"image":1069},"\u002Fblog\u002Fguides\u002Freddit-scraper-returns-different-posts","Two Reddit Scrapers, One Query, Four Posts in Common","The complaint is that a Reddit scraper returns irrelevant posts. We measured it. Results were relevant, and two providers disagreed on six of ten.",[1612,1613,1614,1615,1557],"reddit scraper","reddit api","social listening","apify","\u002Fimg\u002Fblog\u002Freddit-scraper-returns-different-posts.png","\u002Fimg\u002Fblog\u002Freddit-scraper-returns-different-posts-card.png",{"path":1619,"title":1620,"description":1621,"publishedAt":1575,"author":6,"category":1106,"tags":1622,"readTime":1075,"cover":1628,"listingCover":1629,"image":1069},"\u002Fblog\u002Fguides\u002Fsalary-data-api-three-sources","Salary Data API: Three Sources, Three Different Numbers","Indeed says the median software engineer earns 136k. Levels.fyi says 298k. Both are right, and the field that explains it is the sample count.",[1623,1624,1625,1626,1627],"salary data api","compensation data","levels.fyi","h1b salaries","market data","\u002Fimg\u002Fblog\u002Fsalary-data-api-three-sources.png","\u002Fimg\u002Fblog\u002Fsalary-data-api-three-sources-card.png",{"path":1631,"title":1632,"description":1633,"publishedAt":1575,"author":6,"category":1184,"tags":1634,"readTime":1099,"cover":1638,"listingCover":1639,"image":1069},"\u002Fblog\u002Fguides\u002Ftelegram-channel-data-without-a-bot-token","Telegram Channel Data Without a Bot Token","Telegram publishes view counts no other platform gives you. They arrive as 926K and 24.2M, already rounded, and they are lifetime totals not velocity.",[1635,1636,1614,1637,1557],"telegram api","telegram scraping","channel data","\u002Fimg\u002Fblog\u002Ftelegram-channel-data-without-a-bot-token.png","\u002Fimg\u002Fblog\u002Ftelegram-channel-data-without-a-bot-token-card.png",{"path":1641,"title":1642,"description":1643,"publishedAt":1644,"author":6,"category":1092,"tags":1645,"readTime":1075,"cover":1651,"listingCover":1652,"image":1069},"\u002Fblog\u002Fguides\u002Finsider-trading-api-form-4-data","Insider Trading API: What Form 4 Data Actually Tells You","A director sold 1,848,501 shares for $410,843,671. The filing arrived two days later, and that lag is the most important field in the response.","2026-09-07",[1646,1647,1648,1649,1650],"insider trading api","sec form 4","insider transactions","financial data","market signals","\u002Fimg\u002Fblog\u002Finsider-trading-api-form-4-data.png","\u002Fimg\u002Fblog\u002Finsider-trading-api-form-4-data-card.png",{"path":1654,"title":1655,"description":1656,"publishedAt":1644,"author":6,"category":1106,"tags":1657,"readTime":1177,"cover":1663,"listingCover":1664,"image":1069},"\u002Fblog\u002Fguides\u002Fis-there-an-api-for-pitchbook-data","PitchBook Pricing and Cost: Is There an API Instead?","PitchBook publishes no price and no public API. Free alternatives return which rounds a company raised and what type, not how much, when or the valuation.",[1658,1659,1660,1423,1661,1662],"pitchbook pricing","pitchbook cost","pitchbook api","private company data","venture data","\u002Fimg\u002Fblog\u002Fis-there-an-api-for-pitchbook-data.png","\u002Fimg\u002Fblog\u002Fis-there-an-api-for-pitchbook-data-card.png",{"path":1666,"title":1667,"description":1668,"publishedAt":1644,"author":6,"category":1308,"tags":1669,"readTime":1075,"cover":1673,"listingCover":1674,"image":1069},"\u002Fblog\u002Fguides\u002Fmonitor-competitor-prices-automatically","How to Monitor Competitor Prices Automatically","Two sources agreed the price was 54.99. One also said the list price was 79.99. Store only the first and you cannot tell a price cut from a discount.",[1670,1671,1604,1602,1672],"competitor price monitoring","price tracking","list price","\u002Fimg\u002Fblog\u002Fmonitor-competitor-prices-automatically.png","\u002Fimg\u002Fblog\u002Fmonitor-competitor-prices-automatically-card.png",{"path":1676,"title":1677,"description":1678,"publishedAt":1644,"author":6,"category":1092,"tags":1679,"readTime":1177,"cover":1685,"listingCover":1686,"image":1069},"\u002Fblog\u002Fguides\u002Fprediction-market-data-api-kalshi-polymarket","Kalshi API and Polymarket: Matching Prediction Markets by the Call","Two Kalshi routes returned 15 and 25 fields for one venue, with no tickers in common. The cross-venue match endpoint returned 20 pairs at 100 percent.",[1680,1681,1682,1683,1627,1684],"kalshi api","polymarket api","prediction markets","prediction market data","cross-venue","\u002Fimg\u002Fblog\u002Fprediction-market-data-api-kalshi-polymarket.png","\u002Fimg\u002Fblog\u002Fprediction-market-data-api-kalshi-polymarket-card.png",{"path":1688,"title":1689,"description":1690,"publishedAt":1644,"author":6,"category":1092,"tags":1691,"readTime":1075,"cover":1697,"listingCover":1698,"image":1069},"\u002Fblog\u002Fguides\u002Fsoftware-review-data-api-g2-capterra","Software Review Data: What a G2 or Capterra API Returns","Capterra returned 25 reviews split into pros and cons. A niche product returned zero. G2 was blocked on both days we tried, five days apart.",[1692,1693,1694,1695,1696],"software review data","g2 api","capterra api","review mining","competitive research","\u002Fimg\u002Fblog\u002Fsoftware-review-data-api-g2-capterra.png","\u002Fimg\u002Fblog\u002Fsoftware-review-data-api-g2-capterra-card.png",{"path":1700,"title":1701,"description":1702,"publishedAt":1703,"author":6,"category":1184,"tags":1704,"readTime":1075,"cover":1709,"listingCover":1710,"image":1069},"\u002Fblog\u002Fguides\u002Fare-instagram-engagement-trackers-accurate","Are Instagram Follower and Engagement Trackers Accurate?","Two providers disagreed by 2 followers out of 104 million. The same 12 posts gave 0.096% or 0.334% engagement depending on one undocumented choice.","2026-09-03",[1705,1706,1707,1521,1708],"instagram engagement rate","follower count accuracy","influencer vetting","instagram api","\u002Fimg\u002Fblog\u002Fare-instagram-engagement-trackers-accurate.png","\u002Fimg\u002Fblog\u002Fare-instagram-engagement-trackers-accurate-card.png",{"path":1712,"title":1713,"description":1714,"publishedAt":1703,"author":6,"category":1106,"tags":1715,"readTime":1075,"cover":1720,"listingCover":1721,"image":1069},"\u002Fblog\u002Fguides\u002Fbounce-rates-and-spam-traps-when-validating-email","Bounce Rates and Spam Traps When Validating an Email List","A verifier returned twelve fields for three addresses. None of them was spam_trap, and none could be. Here is what verification fixes and what it cannot.",[1716,1717,1534,1718,1719],"spam traps","email bounce rate","deliverability","list hygiene","\u002Fimg\u002Fblog\u002Fbounce-rates-and-spam-traps-when-validating-email.png","\u002Fimg\u002Fblog\u002Fbounce-rates-and-spam-traps-when-validating-email-card.png",{"path":1723,"title":1724,"description":1725,"publishedAt":1703,"author":6,"category":1092,"tags":1726,"readTime":1177,"cover":1732,"listingCover":1733,"image":1069},"\u002Fblog\u002Fguides\u002Frate-limit-calls-to-a-third-party-api","How to Rate Limit Calls to a Third Party API","Six real failure responses collected over three days. Only one of them meant slow down, and telling them apart matters more than any backoff algorithm.",[1727,1728,1729,1730,1731],"rate limiting","retry backoff","429 too many requests","api reliability","concurrency","\u002Fimg\u002Fblog\u002Frate-limit-calls-to-a-third-party-api.png","\u002Fimg\u002Fblog\u002Frate-limit-calls-to-a-third-party-api-card.png",{"path":1735,"title":1736,"description":1737,"publishedAt":1703,"author":6,"category":1106,"tags":1738,"readTime":1075,"cover":1744,"listingCover":1745,"image":1069},"\u002Fblog\u002Fguides\u002Fscrape-linkedin-without-getting-banned","How to Scrape LinkedIn Without Getting Your Account Banned","The ban risk comes from tools that drive your logged-in session. Three endpoints returned profile data from a URL alone, with no cookie involved.",[1739,1740,1741,1742,1743],"linkedin scraping","account ban","li_at cookie","sales navigator","linkedin api","\u002Fimg\u002Fblog\u002Fscrape-linkedin-without-getting-banned.png","\u002Fimg\u002Fblog\u002Fscrape-linkedin-without-getting-banned-card.png",{"path":1747,"title":1748,"description":1749,"publishedAt":1703,"author":6,"category":1092,"tags":1750,"readTime":1177,"cover":1756,"listingCover":1757,"image":1069},"\u002Fblog\u002Fguides\u002Fstop-depending-on-one-scraping-vendor","How to Stop Depending on One Scraping Vendor","Ten provider failures in three days, all while doing other work. None was a bad vendor. The defence is a second route you can reach without a rewrite.",[1751,1752,1753,1754,1755],"vendor lock-in","scraping reliability","actor removed","failover","data pipeline","\u002Fimg\u002Fblog\u002Fstop-depending-on-one-scraping-vendor.png","\u002Fimg\u002Fblog\u002Fstop-depending-on-one-scraping-vendor-card.png",{"path":1759,"title":1760,"description":1761,"publishedAt":1762,"author":6,"category":1092,"tags":1763,"readTime":1177,"cover":1769,"listingCover":1770,"image":1069},"\u002Fblog\u002Fguides\u002Fextract-pricing-and-free-trial-information","Extract Pricing and Free Trial Information From Websites","One endpoint returned four plans, numeric amounts and a free_trial_available flag. It also disclosed that a model read the page, which changes everything.","2026-09-02",[1764,1765,1766,1767,1768],"pricing page extraction","competitor pricing","free trial data","saas pricing","llm extraction","\u002Fimg\u002Fblog\u002Fextract-pricing-and-free-trial-information.png","\u002Fimg\u002Fblog\u002Fextract-pricing-and-free-trial-information-card.png",{"path":1772,"title":1773,"description":1774,"publishedAt":1762,"author":6,"category":1064,"tags":1775,"readTime":1177,"cover":1782,"listingCover":1783,"image":1069},"\u002Fblog\u002Fguides\u002Ffind-all-urls-on-a-domain","Find All Pages on a Domain: Sitemap, Map or Crawl","One map call returned 1,139 pages for a site whose sitemap declared 1,113. Three of four sitemaps returned an index and no count at all.",[1776,1777,1778,1779,1780,1781],"find all pages on a domain","find all urls on a domain","sitemap parser","site map api","url discovery","seo audit","\u002Fimg\u002Fblog\u002Ffind-all-urls-on-a-domain.png","\u002Fimg\u002Fblog\u002Ffind-all-urls-on-a-domain-card.png",{"path":1785,"title":1786,"description":1787,"publishedAt":1762,"author":6,"category":1120,"tags":1788,"readTime":1177,"cover":1794,"listingCover":1795,"image":1069},"\u002Fblog\u002Fguides\u002Fmls-real-estate-api-without-mls-access","MLS Real Estate API: What You Get When You Cannot Get MLS","MLS access needs a licence and a broker. Two public sources returned 40 listings and 48 agents with phones. What they carry, and what they do not.",[1789,1790,1791,1792,1793],"mls real estate api","idx feed","property data","real estate agents","listing data","\u002Fimg\u002Fblog\u002Fmls-real-estate-api-without-mls-access.png","\u002Fimg\u002Fblog\u002Fmls-real-estate-api-without-mls-access-card.png",{"path":1797,"title":1798,"description":1799,"publishedAt":1762,"author":6,"category":1092,"tags":1800,"readTime":1075,"cover":1805,"listingCover":1806,"image":1069},"\u002Fblog\u002Fguides\u002Fsimply-wall-st-graphql-api-fundamentals","Simply Wall St GraphQL API: What Returns Fundamentals Instead","There is no public Simply Wall St GraphQL API. An endpoint that does return fundamentals gave 26 line items across six periods, and named its own source.",[1801,1802,1803,1649,1804],"simply wall st api","stock fundamentals api","income statement api","graphql","\u002Fimg\u002Fblog\u002Fsimply-wall-st-graphql-api-fundamentals.png","\u002Fimg\u002Fblog\u002Fsimply-wall-st-graphql-api-fundamentals-card.png",{"path":1808,"title":1809,"description":1810,"publishedAt":1811,"author":6,"category":1308,"tags":1812,"readTime":1075,"cover":1817,"listingCover":1818,"image":1069},"\u002Fblog\u002Fguides\u002Famazon-asin-scraper-when-it-returns-nothing","Amazon ASIN Scraper: What to Do When It Returns Nothing","A purpose-built ASIN endpoint returned 26 fields and zero values for two valid ASINs. A generic extractor read the same page fine. Assert on values.","2026-09-01",[1813,1814,1602,1815,1816],"amazon asin scraper","asin lookup","product data","silent failure","\u002Fimg\u002Fblog\u002Famazon-asin-scraper-when-it-returns-nothing.png","\u002Fimg\u002Fblog\u002Famazon-asin-scraper-when-it-returns-nothing-card.png",{"path":1820,"title":1821,"description":1822,"publishedAt":1811,"author":6,"category":1092,"tags":1823,"readTime":1177,"cover":1829,"listingCover":1830,"image":1069},"\u002Fblog\u002Fguides\u002Fdatacenter-proxies-vs-residential","Datacenter Proxies vs Residential: What Each One Actually Fixes","A real residential IP with a real Chrome user agent still got 403 from two sites. The IP type was not the constraint. What each type does fix.",[1824,1825,1826,1827,1828],"datacenter proxies vs residential","proxy types","rotating proxies","web scraping","geolocation","\u002Fimg\u002Fblog\u002Fdatacenter-proxies-vs-residential.png","\u002Fimg\u002Fblog\u002Fdatacenter-proxies-vs-residential-card.png",{"path":1832,"title":1833,"description":1834,"publishedAt":1811,"author":6,"category":1106,"tags":1835,"readTime":1177,"cover":1842,"listingCover":1843,"image":1069},"\u002Fblog\u002Fguides\u002Fextracting-leadership-contact-information","Email Extractor for a Domain: 1,333 Addresses, and What a Score Hides","One domain returned 1,333 known addresses. Two rows both scored 99, one corroborated eight times and one just once. The score hides that.",[1836,1837,1838,1839,1840,1841],"email extractor","email extractor extension","linkedin email finder","leadership contact information","decision maker search","hunter io","\u002Fimg\u002Fblog\u002Fextracting-leadership-contact-information.png","\u002Fimg\u002Fblog\u002Fextracting-leadership-contact-information-card.png",{"path":1845,"title":1846,"description":1847,"publishedAt":1811,"author":6,"category":1092,"tags":1848,"readTime":1177,"cover":1854,"listingCover":1855,"image":1069},"\u002Fblog\u002Fguides\u002Fpuppeteer-extra-plugin-stealth-release-date","Puppeteer Extra Plugin Stealth: Check the Release Date First","The stealth plugin last shipped in March 2023. Puppeteer shipped last week. What that gap means, and what scrapy-playwright does differently.",[1849,1850,1851,1852,1853],"puppeteer-extra-plugin-stealth","scrapy-playwright","headless browser","bot detection","browser automation","\u002Fimg\u002Fblog\u002Fpuppeteer-extra-plugin-stealth-release-date.png","\u002Fimg\u002Fblog\u002Fpuppeteer-extra-plugin-stealth-release-date-card.png",{"path":1857,"title":1858,"description":1859,"publishedAt":1860,"author":6,"category":1064,"tags":1861,"readTime":1099,"cover":1867,"listingCover":1868,"image":1069},"\u002Fblog\u002Fguides\u002Fc-sharp-web-scraping-without-the-stack","C# Web Scraping: Keep HtmlAgilityPack, Drop the Fetch","The .NET parsing stack was never the problem. What breaks a C# scraper is the request leaving your server, and that half is not a language question.","2026-08-31",[1862,1863,1864,1865,1866],"c# web scraping","htmlagilitypack","dotnet","web extraction","playwright","\u002Fimg\u002Fblog\u002Fc-sharp-web-scraping-without-the-stack.png","\u002Fimg\u002Fblog\u002Fc-sharp-web-scraping-without-the-stack-card.png",{"path":1870,"title":1871,"description":1872,"publishedAt":1860,"author":6,"category":1106,"tags":1873,"readTime":1075,"cover":1879,"listingCover":1880,"image":1069},"\u002Fblog\u002Fguides\u002Flinkedin-post-scraper-what-comes-back","LinkedIn Post Scraper: What Actually Comes Back","Posts, reactions and comments are three billing decisions in one call. A live run shows the fields, and the two flags that quietly multiply the bill.",[1874,1875,1876,1877,1878],"linkedin post scraper","linkedin company scraper","social selling","harvestapi","linkedin data","\u002Fimg\u002Fblog\u002Flinkedin-post-scraper-what-comes-back.png","\u002Fimg\u002Fblog\u002Flinkedin-post-scraper-what-comes-back-card.png",{"path":1882,"title":1883,"description":1884,"publishedAt":1860,"author":6,"category":1064,"tags":1885,"readTime":1075,"cover":1890,"listingCover":1891,"image":1069},"\u002Fblog\u002Fguides\u002Fweb-scraping-vs-api-wrong-question","Web Scraping vs API: You Are Choosing a Maintainer","Both end in an HTTP request returning the same facts. What differs is who owns the parser when the page changes, and whether anyone promised you anything.",[1886,1887,1888,1889,1567],"web scraping vs api","data acquisition","buy vs build","scraping api","\u002Fimg\u002Fblog\u002Fweb-scraping-vs-api-wrong-question.png","\u002Fimg\u002Fblog\u002Fweb-scraping-vs-api-wrong-question-card.png",{"path":1893,"title":1894,"description":1895,"publishedAt":1860,"author":6,"category":1064,"tags":1896,"readTime":1075,"cover":1901,"listingCover":1902,"image":1069},"\u002Fblog\u002Fguides\u002Fxpath-cheat-sheet-for-scraping","XPath Cheat Sheet: The Dozen Expressions That Survive","Most XPath references list what the spec allows. This lists what still matches after a redesign, ranked by how long each anchor survives.",[1897,1898,1899,1865,1900],"xpath cheat sheet","xpath","css selectors","scraping","\u002Fimg\u002Fblog\u002Fxpath-cheat-sheet-for-scraping.png","\u002Fimg\u002Fblog\u002Fxpath-cheat-sheet-for-scraping-card.png",{"path":1904,"title":1905,"description":1906,"publishedAt":1860,"author":6,"category":1120,"tags":1907,"readTime":1913,"cover":1914,"listingCover":1915,"image":1069},"\u002Fblog\u002Fguides\u002Fzillow-scraper-what-the-listing-carries","Zillow API: One Search Returned 41, 502 and 874 at the Same Time","Three count fields in one response, all different. Then one filter took the same search from 874 results to 4. Which number to read, and which to ignore.",[1908,1909,1910,1911,1912],"zillow api","zillow scraper","property data api","zestimate","real estate api","13 min","\u002Fimg\u002Fblog\u002Fzillow-scraper-what-the-listing-carries.png","\u002Fimg\u002Fblog\u002Fzillow-scraper-what-the-listing-carries-card.png",{"path":1917,"title":1918,"description":1919,"publishedAt":1920,"author":6,"category":1586,"tags":1921,"readTime":1075,"cover":1924,"listingCover":1925,"image":1069},"\u002Fblog\u002Fguides\u002Fad-agent-auction-outside-the-account","Soku Runs the Campaign. The Auction Sits Outside It.","One small advertiser shares a single keyword with a bidder running 198,105 of them. An ad agent optimising inside its own account cannot see that.","2026-08-30",[1591,1922,1923,1567],"ad agents","competitor intelligence","\u002Fimg\u002Fblog\u002Fad-agent-auction-outside-the-account.png","\u002Fimg\u002Fblog\u002Fad-agent-auction-outside-the-account-card.png",{"path":1927,"title":1928,"description":1929,"publishedAt":1920,"author":6,"category":1064,"tags":1930,"readTime":1075,"cover":1935,"listingCover":1936,"image":1069},"\u002Fblog\u002Fguides\u002Fbing-serp-tracker-after-the-api","Bing SERP Tracker: What to Use After the API Retired","Microsoft retired the Search API, so a Bing position now comes from three places. One is free and first-party, and most articles do not mention it.",[1931,1932,1275,1933,1934],"bing serp tracker","bing rank tracking","bing webmaster tools","rank tracking","\u002Fimg\u002Fblog\u002Fbing-serp-tracker-after-the-api.png","\u002Fimg\u002Fblog\u002Fbing-serp-tracker-after-the-api-card.png",{"path":1938,"title":1939,"description":1940,"publishedAt":1920,"author":6,"category":1092,"tags":1941,"readTime":1075,"cover":1945,"listingCover":1946,"image":1069},"\u002Fblog\u002Fguides\u002Fbrowser-automation-without-an-api","Browser Automation When the Site Has No API","Check the premise first. The site owner has no API, but somebody may already have an endpoint for it, and a browser is the most expensive answer available.",[1853,1942,1943,1944,1567],"antidetect browser","nodriver","no api","\u002Fimg\u002Fblog\u002Fbrowser-automation-without-an-api.png","\u002Fimg\u002Fblog\u002Fbrowser-automation-without-an-api-card.png",{"path":1948,"title":1949,"description":1950,"publishedAt":1920,"author":6,"category":1106,"tags":1951,"readTime":1075,"cover":1957,"listingCover":1958,"image":1069},"\u002Fblog\u002Fguides\u002Fjob-scraping-software-per-board","Job Scraping Software: There Is No Single Job Feed","Postings are fragmented across boards by design, and no endpoint covers all of them. Pick per board, and the schema differences are the real work.",[1952,1953,1954,1955,1956],"job scraping software","indeed job scraper","jobspy","job postings api","hiring signals","\u002Fimg\u002Fblog\u002Fjob-scraping-software-per-board.png","\u002Fimg\u002Fblog\u002Fjob-scraping-software-per-board-card.png",{"path":1960,"title":1961,"description":1962,"publishedAt":1920,"author":6,"category":1106,"tags":1963,"readTime":1075,"cover":1967,"listingCover":1968,"image":1069},"\u002Fblog\u002Fguides\u002Fmid-form-enrichment-fewer-questions","TinyCommand Asks One Question. The API Fills Twenty-Six.","One domain went in. Twenty-six fields came back, including a 43-item tech stack. Every one of them is a question the form no longer has to ask.",[1964,1965,1966,1250],"lead enrichment","forms","progressive profiling","\u002Fimg\u002Fblog\u002Fmid-form-enrichment-fewer-questions.png","\u002Fimg\u002Fblog\u002Fmid-form-enrichment-fewer-questions-card.png",{"path":1970,"title":1971,"description":1972,"publishedAt":1920,"author":6,"category":1064,"tags":1973,"readTime":1099,"cover":1978,"listingCover":1979,"image":1069},"\u002Fblog\u002Fguides\u002Fnode-unblocker-what-it-does-not-do","Node Unblocker: What It Does, and What It Does Not","A proxy you host still leaves from your address. It changes the code path, not what refused you. Its own README lists the sites it cannot handle.",[1974,1975,1976,1977,1865],"node unblocker","web proxy","blocked scraper","nodejs","\u002Fimg\u002Fblog\u002Fnode-unblocker-what-it-does-not-do.png","\u002Fimg\u002Fblog\u002Fnode-unblocker-what-it-does-not-do-card.png",{"path":1981,"title":1982,"description":1983,"publishedAt":1920,"author":6,"category":1064,"tags":1984,"readTime":1177,"cover":1989,"listingCover":1990,"image":1069},"\u002Fblog\u002Fguides\u002Fserp-history-track-the-shape","AI Visibility Tracking: Record the Page's Shape, Not Just Your Rank","A flat position line hides the month an AI block appeared and took the clicks. We have 45 keywords on page one and one click. Here is what to record.",[1985,1986,1987,1988,1934],"ai visibility tracking","serp history","ai overview","paa seo","\u002Fimg\u002Fblog\u002Fserp-history-track-the-shape.png","\u002Fimg\u002Fblog\u002Fserp-history-track-the-shape-card.png",{"path":1992,"title":1993,"description":1994,"publishedAt":1995,"author":6,"category":1064,"tags":1996,"readTime":1177,"cover":2002,"listingCover":2003,"image":1069},"\u002Fblog\u002Fguides\u002Fbrave-search-api-what-you-can-keep","Brave Search API Key: How to Get One, and What You May Keep","Getting the key is the easy half. One clause decides whether it fits: result storage is granted per plan, and most builds need to store. Plus a free route.","2026-08-27",[1997,1998,1999,2000,2001,1079],"brave search api key","brave api key","how to get brave api key","brave search api","independent index","\u002Fimg\u002Fblog\u002Fbrave-search-api-what-you-can-keep.png","\u002Fimg\u002Fblog\u002Fbrave-search-api-what-you-can-keep-card.png",{"path":2005,"title":2006,"description":2007,"publishedAt":1995,"author":6,"category":1064,"tags":2008,"readTime":1075,"cover":2012,"listingCover":2013,"image":1069},"\u002Fblog\u002Fguides\u002Fgoogle-search-url-parameters","Google Search URL Parameters: What Still Works in 2026","Google silently killed num in September 2025. Here is the current list, which ones are official, and which are community guesses you should not depend on.",[2009,1275,1934,2010,2011],"google search url parameters","udm","search operators","\u002Fimg\u002Fblog\u002Fgoogle-search-url-parameters.png","\u002Fimg\u002Fblog\u002Fgoogle-search-url-parameters-card.png",{"path":2015,"title":2016,"description":2017,"publishedAt":1995,"author":6,"category":1184,"tags":2018,"readTime":1177,"cover":2023,"listingCover":2024,"image":1069},"\u002Fblog\u002Fguides\u002Finstagram-api-which-one-in-2026","Instagram Profile Download and Search: Which API in 2026?","One profile call returned 69 fields and 5.5 million followers nested inside a wrapper. Four products answer to the name Instagram API.",[2019,2020,2021,1708,2022,1557],"instagram profile download","download instagram profile","instagram profile search","instagram graph api","\u002Fimg\u002Fblog\u002Finstagram-api-which-one-in-2026.png","\u002Fimg\u002Fblog\u002Finstagram-api-which-one-in-2026-card.png",{"path":2026,"title":2027,"description":2028,"publishedAt":1995,"author":6,"category":1184,"tags":2029,"readTime":1075,"cover":2031,"listingCover":2032,"image":1069},"\u002Fblog\u002Fguides\u002Ftiktok-scraper-which-endpoint","TikTok Scraper: Which Endpoint for Which Question?","Profiles, videos, comments and search are four jobs with four cost curves. And the creator arrives nested inside every post, which changes the design.",[2030,1556,1521,1557,1615],"tiktok scraper","\u002Fimg\u002Fblog\u002Ftiktok-scraper-which-endpoint.png","\u002Fimg\u002Fblog\u002Ftiktok-scraper-which-endpoint-card.png",{"path":2034,"title":2035,"description":2036,"publishedAt":1995,"author":6,"category":1064,"tags":2037,"readTime":1075,"cover":2041,"listingCover":2042,"image":1069},"\u002Fblog\u002Fguides\u002Fweb-scraping-services-should-you-hire-it-out","Web Scraping Services: Should You Hire This Out?","A managed service is priced against the value of your data, not the cost of the requests. Three questions decide whether that is a bargain or a markup.",[2038,2039,1888,2040,1889],"web scraping services","managed scraping","data as a service","\u002Fimg\u002Fblog\u002Fweb-scraping-services-should-you-hire-it-out.png","\u002Fimg\u002Fblog\u002Fweb-scraping-services-should-you-hire-it-out-card.png",{"path":2044,"title":2045,"description":2046,"publishedAt":1084,"author":6,"category":1092,"tags":2047,"readTime":1075,"cover":2051,"listingCover":2052,"image":1069},"\u002Fblog\u002Fguides\u002Fbright-data-mcp-what-you-are-installing","Bright Data MCP: What Are You Actually Installing?","Sixty-nine tools arrive in one install, and every definition sits in the model's context before the agent does anything. When that trade is worth it.",[2048,1484,1567,2049,2050],"bright data mcp","web data","tool calling","\u002Fimg\u002Fblog\u002Fbright-data-mcp-what-you-are-installing.png","\u002Fimg\u002Fblog\u002Fbright-data-mcp-what-you-are-installing-card.png",{"path":2054,"title":2055,"description":2056,"publishedAt":1084,"author":6,"category":1106,"tags":2057,"readTime":1075,"cover":2062,"listingCover":2063,"image":1069},"\u002Fblog\u002Fguides\u002Fclay-alternatives-waterfall-enrichment","Clay Alternatives: You Are Buying the Waterfall","Clay's product is not the data, it is the fall-through logic between providers. Once they share a balance, that logic is about twenty lines of code.",[2058,2059,2060,2061],"clay alternatives","waterfall enrichment","lead enrichment api","fall-through logic","\u002Fimg\u002Fblog\u002Fclay-alternatives-waterfall-enrichment.png","\u002Fimg\u002Fblog\u002Fclay-alternatives-waterfall-enrichment-card.png",{"path":2065,"title":2066,"description":2067,"publishedAt":1084,"author":6,"category":1492,"tags":2068,"readTime":1099,"cover":2074,"listingCover":2075,"image":1069},"\u002Fblog\u002Fguides\u002Felevenlabs-alternatives-three-complaints","ElevenLabs Alternatives: Price, Voice, or Plumbing?","Three different complaints hide under one search. Only one of them needs a different provider, and the other two have cheaper fixes on the same one.",[2069,2070,2071,2072,2073],"elevenlabs alternatives","text to speech","tts api","voice ai","generative media","\u002Fimg\u002Fblog\u002Felevenlabs-alternatives-three-complaints.png","\u002Fimg\u002Fblog\u002Felevenlabs-alternatives-three-complaints-card.png",{"path":2077,"title":2078,"description":2079,"publishedAt":1084,"author":6,"category":1064,"tags":2080,"readTime":1075,"cover":2086,"listingCover":2087,"image":1069},"\u002Fblog\u002Fguides\u002Frank-tracking-api-store-the-position","Rank Tracking API: The Hard Part Is Not the Fetch","Getting today's position is one cheap call. The product is yesterday's position, identical parameters, and knowing what moved. Here is the whole shape.",[2081,2082,2083,2084,2085],"rank tracking api","bing rank tracker","serp tracking","seo api","search api","\u002Fimg\u002Fblog\u002Frank-tracking-api-store-the-position.png","\u002Fimg\u002Fblog\u002Frank-tracking-api-store-the-position-card.png",{"path":2089,"title":2090,"description":2091,"publishedAt":1084,"author":6,"category":1492,"tags":2092,"readTime":1075,"cover":2096,"listingCover":2097,"image":1069},"\u002Fblog\u002Fguides\u002Fstreaming-voice-agent-vs-batch-tts","Gradium Holds the Call. A Batch TTS API Renders the Rest.","Ten measured calls to a batch text to speech endpoint. The fastest never beat 3.05 seconds, and the model mattered more than four times the text.",[2093,2094,2095,1567],"text to speech api","voice agents","tts latency","\u002Fimg\u002Fblog\u002Fstreaming-voice-agent-vs-batch-tts.png","\u002Fimg\u002Fblog\u002Fstreaming-voice-agent-vs-batch-tts-card.png",{"path":2099,"title":2100,"description":2101,"publishedAt":2102,"author":6,"category":1106,"tags":2103,"readTime":1075,"cover":2109,"listingCover":2110,"image":1069},"\u002Fblog\u002Fguides\u002Fb2b-data-provider-dataset-vs-api","B2B Data Providers: Buy the Dataset or Call the API?","A dataset and an enrichment API are the same rows with a different contract. What changes is who carries the staleness and who pays for unread records.","2026-08-25",[2104,2105,2106,2107,2108],"b2b data","data provider","company enrichment","zoominfo","people data labs","\u002Fimg\u002Fblog\u002Fb2b-data-provider-dataset-vs-api.png","\u002Fimg\u002Fblog\u002Fb2b-data-provider-dataset-vs-api-card.png",{"path":2112,"title":2113,"description":2114,"publishedAt":2102,"author":6,"category":1064,"tags":2115,"readTime":1075,"cover":2119,"listingCover":2120,"image":1069},"\u002Fblog\u002Fguides\u002Fcrawl4ai-when-to-run-your-own-crawler","Crawl4AI: When Should You Run Your Own Crawler?","Crawl4AI does the parsing for free. The bill is the fleet you run around it. A decision rule for when to self-host a crawler and when to call one.",[2116,2117,1079,1083,2118],"crawl4ai","web crawler","open source","\u002Fimg\u002Fblog\u002Fcrawl4ai-when-to-run-your-own-crawler.png","\u002Fimg\u002Fblog\u002Fcrawl4ai-when-to-run-your-own-crawler-card.png",{"path":2122,"title":2123,"description":2124,"publishedAt":2102,"author":6,"category":1064,"tags":2125,"readTime":1177,"cover":2131,"listingCover":2132,"image":1069},"\u002Fblog\u002Fguides\u002Fis-web-scraping-legal-2026","Is Web Scraping Legal? What the Rulings Actually Say","Four US cases decided the modern law of scraping and none of them turned on scraping. They turned on login, publicness, copying and circumvention.",[2126,2127,2128,2129,2130],"web scraping legal","hiq linkedin","bright data","cfaa","dmca","\u002Fimg\u002Fblog\u002Fis-web-scraping-legal-2026.png","\u002Fimg\u002Fblog\u002Fis-web-scraping-legal-2026-card.png",{"path":2134,"title":2135,"description":2136,"publishedAt":2102,"author":6,"category":1064,"tags":2137,"readTime":1075,"cover":2141,"listingCover":2142,"image":1069},"\u002Fblog\u002Fguides\u002Fweb-scraping-news-articles","Web Scraping News Articles: Three Jobs, Three Endpoints","Scraping news is three jobs under one name. Find articles, read one article, watch one company. Each needs a different call and a different budget.",[2138,2139,1079,2140,1566],"news scraping","news api","media monitoring","\u002Fimg\u002Fblog\u002Fweb-scraping-news-articles.png","\u002Fimg\u002Fblog\u002Fweb-scraping-news-articles-card.png",{"path":2144,"title":2145,"description":2146,"publishedAt":2102,"author":6,"category":1184,"tags":2147,"readTime":1075,"cover":2150,"listingCover":2151,"image":1069},"\u002Fblog\u002Fguides\u002Fyoutube-scraper-past-the-quota","YouTube Scraper: What the Data API Will Not Give You","The official API allows 100 searches a day and no money raises it. That ceiling, not price, is why YouTube scrapers exist. Here is what each one returns.",[2148,2149,1521,1615,1557],"youtube scraper","youtube data api","\u002Fimg\u002Fblog\u002Fyoutube-scraper-past-the-quota.png","\u002Fimg\u002Fblog\u002Fyoutube-scraper-past-the-quota-card.png",{"path":2153,"title":2154,"description":2155,"publishedAt":2156,"author":6,"category":1092,"tags":2157,"readTime":1177,"cover":2160,"listingCover":2161,"image":1069},"\u002Fblog\u002Fguides\u002Fagentic-browser-when-your-agent-needs-one","Agentic Browser: When Your Agent Actually Needs One","Agentic browser names two products: one you sit in front of and one your agent drives. Most jobs need neither. The test is whether the task carries state.","2026-08-21",[2158,1853,1567,1866,2159],"agentic browser","browserbase","\u002Fimg\u002Fblog\u002Fagentic-browser-when-your-agent-needs-one.png","\u002Fimg\u002Fblog\u002Fagentic-browser-when-your-agent-needs-one-card.png",{"path":2163,"title":2164,"description":2165,"publishedAt":2156,"author":6,"category":1064,"tags":2166,"readTime":1177,"cover":2170,"listingCover":2171,"image":1069},"\u002Fblog\u002Fguides\u002Fdo-ai-agents-need-a-rotating-proxy","Do You Still Need a Rotating Proxy in 2026?","A rotating proxy is an input you buy by the gigabyte and hope works. Most teams wanted the outcome instead. Here is how to tell which one you need.",[2167,2168,1827,1567,2169],"rotating proxy","residential proxy","infrastructure","\u002Fimg\u002Fblog\u002Fdo-ai-agents-need-a-rotating-proxy.png","\u002Fimg\u002Fblog\u002Fdo-ai-agents-need-a-rotating-proxy-card.png",{"path":691,"title":692,"description":2173,"publishedAt":2156,"author":6,"category":1064,"tags":2174,"readTime":1177,"cover":2178,"listingCover":2179,"image":1069},"Python web scraping is two jobs wearing one name. Python is the best tool for one of them and the wrong place to solve the other. Here is the split.",[2175,1827,2176,2177,1567],"python","beautifulsoup","selenium","\u002Fimg\u002Fblog\u002Fpython-web-scraping-without-a-scraper.png","\u002Fimg\u002Fblog\u002Fpython-web-scraping-without-a-scraper-card.png",{"path":2181,"title":2182,"description":2183,"publishedAt":2156,"author":6,"category":1064,"tags":2184,"readTime":1913,"cover":2188,"listingCover":2189,"image":1069},"\u002Fblog\u002Fguides\u002Fserp-api-for-ai-agents","Google SERP Tracker and Analysis: Which SERP API Do You Need?","A SERP API is three products under one name. Tracking needs two rank fields; analysis needs the blocks around them, and one was missing.",[2185,2186,2187,1275,1566,1079],"google serp tracker","google serp analysis","cheapest serp api","\u002Fimg\u002Fblog\u002Fserp-api-for-ai-agents.png","\u002Fimg\u002Fblog\u002Fserp-api-for-ai-agents-card.png",{"path":2191,"title":2192,"description":2193,"publishedAt":2156,"author":6,"category":1064,"tags":2194,"readTime":1177,"cover":2197,"listingCover":2198,"image":1069},"\u002Fblog\u002Fguides\u002Fweb-scraping-tools-which-kind","Web Scraping Tools: Which Kind Do You Actually Need?","Web scraping tools come in four kinds and they differ on who operates them, not on features. Pick the operator first and the shortlist writes itself.",[2195,2196,1889,1567,1615],"web scraping tools","no code scraper","\u002Fimg\u002Fblog\u002Fweb-scraping-tools-which-kind.png","\u002Fimg\u002Fblog\u002Fweb-scraping-tools-which-kind-card.png",{"path":2200,"title":2201,"description":2202,"publishedAt":1074,"author":6,"category":1308,"tags":2203,"readTime":1913,"cover":2208,"listingCover":2209,"image":1069},"\u002Fblog\u002Fguides\u002Fautonomous-store-agent-outside-data","Agentic Commerce: The Store Builds Itself, Not the Market","The hard part moved. Building a store is one sentence now. One call returned 48 competitors and an incumbent with 136,609 reviews.",[2204,2205,2206,2207],"agentic commerce","autonomous store agent","store agent data","market entry research","\u002Fimg\u002Fblog\u002Fautonomous-store-agent-outside-data.png","\u002Fimg\u002Fblog\u002Fautonomous-store-agent-outside-data-card.png",{"path":2211,"title":2212,"description":2213,"publishedAt":1074,"author":6,"category":1064,"tags":2214,"readTime":1177,"cover":2218,"listingCover":2219,"image":1069},"\u002Fblog\u002Fguides\u002Fbing-search-api-retired-alternatives","Bing Search API Retired: What Actually Replaces It","Microsoft turned the Bing Search APIs off in August 2025 and the endpoints now return 410. What a replacement has to do, and which route fits which job.",[2215,936,2216,1079,2217],"search","api","agents","\u002Fimg\u002Fblog\u002Fbing-search-api-retired-alternatives.png","\u002Fimg\u002Fblog\u002Fbing-search-api-retired-alternatives-card.png",{"path":1073,"title":5,"description":1066,"publishedAt":1074,"author":6,"category":1064,"tags":2221,"readTime":1075,"cover":1065,"listingCover":1070,"image":1069},[1079,1080,1081,1082,1083],{"path":2223,"title":2224,"description":2225,"publishedAt":1074,"author":6,"category":1184,"tags":2226,"readTime":1177,"cover":2231,"listingCover":2232,"image":1069},"\u002Fblog\u002Fguides\u002Finstagram-follower-engagement-api","Instagram Follower and Engagement Data: Which API in 2026?","Follower counts are accurate. What trackers build on top is inference. Which endpoints return which fields, and how to tell a real signal from a guess.",[2227,2228,2216,2229,2230],"instagram","social","creators","engagement","\u002Fimg\u002Fblog\u002Finstagram-follower-engagement-api.png","\u002Fimg\u002Fblog\u002Finstagram-follower-engagement-api-card.png",{"path":2234,"title":2235,"description":2236,"publishedAt":1074,"author":6,"category":1092,"tags":2237,"readTime":1177,"cover":2243,"listingCover":2244,"image":1069},"\u002Fblog\u002Fguides\u002Fllm-gateway-vs-mcp-gateway","LLM Gateway vs MCP Gateway: Four Families, One Word","An LLM gateway routes prompts to models. An MCP gateway routes tool calls to vendors. Four gateway families, and which problem each one solves.",[2238,2239,2240,2241,2242],"llm gateway","ai gateway","mcp","agent tools","routing","\u002Fimg\u002Fblog\u002Fllm-gateway-vs-mcp-gateway.png","\u002Fimg\u002Fblog\u002Fllm-gateway-vs-mcp-gateway-card.png",{"path":2246,"title":2247,"description":2248,"publishedAt":1074,"author":6,"category":1092,"tags":2249,"readTime":1177,"cover":2250,"listingCover":2251,"image":1069},"\u002Fblog\u002Fguides\u002Fmcp-vs-api-for-ai-agents","MCP vs API for AI Agents: Who Does the Wrapping?","MCP is not an alternative to APIs, it wraps them. The real question is who does the wrapping, and the answer decides how much code you write.",[2240,2216,2241,1567,2050],"\u002Fimg\u002Fblog\u002Fmcp-vs-api-for-ai-agents.png","\u002Fimg\u002Fblog\u002Fmcp-vs-api-for-ai-agents-card.png",{"path":2253,"title":2254,"description":2255,"publishedAt":1074,"author":6,"category":1092,"tags":2256,"readTime":1075,"cover":2260,"listingCover":2261,"image":1069},"\u002Fblog\u002Fguides\u002Fn8n-data-layer-scraping-enrichment","The Data Layer for n8n: One Key Instead of a Node Per Vendor","n8n workflows rarely break at the logic. They break at the data source. Build the scraping and enrichment layer so a dead provider is a config change.",[2257,2258,2259,1900,2217],"n8n","automation","enrichment","\u002Fimg\u002Fblog\u002Fn8n-data-layer-scraping-enrichment.png","\u002Fimg\u002Fblog\u002Fn8n-data-layer-scraping-enrichment-card.png",{"path":2263,"title":2264,"description":2265,"publishedAt":1074,"author":6,"category":1092,"tags":2266,"readTime":1177,"cover":2268,"listingCover":2269,"image":1069},"\u002Fblog\u002Fguides\u002Fopenrouter-alternatives-after-stripe","OpenRouter Alternatives in 2026: Models, and Then Tools","Stripe agreed to buy OpenRouter for over $7 billion. The real alternatives on the model side, and the routing layer nobody is selling yet.",[2267,2238,2241,2240,2242],"openrouter","\u002Fimg\u002Fblog\u002Fopenrouter-alternatives-after-stripe.png","\u002Fimg\u002Fblog\u002Fopenrouter-alternatives-after-stripe-card.png",{"path":2271,"title":2272,"description":2273,"publishedAt":1074,"author":6,"category":1092,"tags":2274,"readTime":1177,"cover":2278,"listingCover":2279,"image":1069},"\u002Fblog\u002Fguides\u002Fopenrouter-mcp-server-tool-layer","OpenRouter API Key: What It Reaches, and What It Does Not","An OpenRouter key buys models. It does not buy tools. Here is exactly where that line falls, and how to give the same agent both without a second account.",[2275,2267,2276,2050,2277],"openrouter api key","model routing","claude code","\u002Fimg\u002Fblog\u002Fopenrouter-mcp-server-tool-layer.png","\u002Fimg\u002Fblog\u002Fopenrouter-mcp-server-tool-layer-card.png",{"path":2281,"title":2282,"description":2283,"publishedAt":1074,"author":6,"category":1106,"tags":2284,"readTime":1177,"cover":2287,"listingCover":2288,"image":1069},"\u002Fblog\u002Fguides\u002Fprompt-instead-of-filters-people-search","A Wrong Enum Returns Zero: Ploid and Prompt-Shaped APIs","A wrong enum returned 0 rows. A wrong range format returned 160,884. Neither raised an error. Why prompt-shaped APIs like Ploid exist, and what they cost.",[2285,2286,1590,1567],"people search api","agentic api","\u002Fimg\u002Fblog\u002Fprompt-instead-of-filters-people-search.png","\u002Fimg\u002Fblog\u002Fprompt-instead-of-filters-people-search-card.png",{"path":2290,"title":2291,"description":2292,"publishedAt":1074,"author":6,"category":1106,"tags":2293,"readTime":1177,"cover":2297,"listingCover":2298,"image":1069},"\u002Fblog\u002Fguides\u002Fproxycurl-shutdown-linkedin-data-alternatives","Proxycurl Shut Down: Where LinkedIn Data Goes Now","Proxycurl closed in July 2025 to settle with LinkedIn. What a replacement has to return, which endpoints cover which half, and the risk to read first.",[2294,2259,2216,2295,2296],"linkedin","sales","leads","\u002Fimg\u002Fblog\u002Fproxycurl-shutdown-linkedin-data-alternatives.png","\u002Fimg\u002Fblog\u002Fproxycurl-shutdown-linkedin-data-alternatives-card.png",{"path":2300,"title":2301,"description":2302,"publishedAt":1074,"author":6,"category":1184,"tags":2303,"readTime":1177,"cover":2308,"listingCover":2309,"image":1069},"\u002Fblog\u002Fguides\u002Ftiktok-data-behind-an-ai-video-ad","Oumomo Generates the Ad. TikTok Data Decides Which One.","An AI video tool makes whatever you ask. One TikTok search call returned twenty videos with a fifteenfold spread in likes, and that spread is the brief.",[2304,2305,2306,2307],"tiktok data api","ai video ads","tiktok shop","creative research","\u002Fimg\u002Fblog\u002Ftiktok-data-behind-an-ai-video-ad.png","\u002Fimg\u002Fblog\u002Ftiktok-data-behind-an-ai-video-ad-card.png",{"path":2311,"title":2312,"description":2313,"publishedAt":1074,"author":6,"category":1184,"tags":2314,"readTime":1177,"cover":2315,"listingCover":2316,"image":1069},"\u002Fblog\u002Fguides\u002Ftiktok-scraper-catalogue-audit","Oumomo Remakes the Winner. A TikTok Scraper Tells You Which.","Fourteen videos, twenty one days, and one was 65 percent of the account. You cannot see that in the clips. One call puts the catalogue in a table.",[2030,1274,2306,2305],"\u002Fimg\u002Fblog\u002Ftiktok-scraper-catalogue-audit.png","\u002Fimg\u002Fblog\u002Ftiktok-scraper-catalogue-audit-card.png",{"path":2318,"title":2319,"description":2320,"publishedAt":1074,"author":6,"category":1092,"tags":2321,"readTime":1913,"cover":2323,"listingCover":2324,"image":1069},"\u002Fblog\u002Fguides\u002Fwhat-is-an-mcp-gateway","What Is an MCP Gateway? Two Products, One Name","An MCP gateway is one endpoint that fronts many tools. Two very different products share the name, and the difference decides which one you need.",[2240,2322,2241,1567,2242],"mcp gateway","\u002Fimg\u002Fblog\u002Fwhat-is-an-mcp-gateway.png","\u002Fimg\u002Fblog\u002Fwhat-is-an-mcp-gateway-card.png",{"path":2326,"title":2327,"description":2328,"publishedAt":1074,"author":6,"category":1184,"tags":2329,"readTime":1177,"cover":2332,"listingCover":2333,"image":1069},"\u002Fblog\u002Fguides\u002Fx-intent-search-social-selling","Monid Finds the Tweet. Volumn Tells You Who Sent It.","A tightened X search returned 19 of 20 keyword-relevant tweets. Only 2 were someone actually asking. Keyword match is not intent, and that gap is the work.",[2330,2331,1876,1239],"twitter scraper","x api","\u002Fimg\u002Fblog\u002Fx-intent-search-social-selling.png","\u002Fimg\u002Fblog\u002Fx-intent-search-social-selling-card.png",{"path":2335,"title":2336,"description":2337,"publishedAt":2338,"author":6,"category":1064,"tags":2339,"readTime":1075,"cover":2343,"listingCover":2344,"image":1069},"\u002Fblog\u002Fguides\u002Fapify-alternatives","Apify Alternatives: You Probably Want a Different Way to Buy It","Every AI answer to this question names Apify. We ran an Apify actor without an Apify account to show what the real alternative is.","2026-08-19",[2340,1827,2341,2342],"apify alternatives","actors","pay per call","\u002Fimg\u002Fblog\u002Fapify-alternatives.png","\u002Fimg\u002Fblog\u002Fapify-alternatives-card.png",{"path":2346,"title":2347,"description":2348,"publishedAt":2338,"author":6,"category":1308,"tags":2349,"readTime":1075,"cover":2354,"listingCover":2355,"image":1069},"\u002Fblog\u002Fguides\u002Fconnect-claude-to-amazon-search-data","Amazon Search API: What One Keyword Query Returns","Amazon has no keyword API for sellers. What a search endpoint actually returns, how to hand it to an agent, and which route fits which job.",[2350,2351,2352,2353],"amazon search api","amazon keyword data","amazon product api","seller research","\u002Fimg\u002Fblog\u002Fconnect-claude-to-amazon-search-data.png","\u002Fimg\u002Fblog\u002Fconnect-claude-to-amazon-search-data-card.png",{"path":2357,"title":2358,"description":2359,"publishedAt":2338,"author":6,"category":1106,"tags":2360,"readTime":1075,"cover":2365,"listingCover":2366,"image":1069},"\u002Fblog\u002Fguides\u002Fenrich-a-list-from-email-addresses","Email Enrichment: Four Addresses In, One Person Out","We enriched four real email addresses and one returned a person. The misses billed nothing. What that changes about how you plan a batch.",[2361,2362,2363,2364],"email enrichment","email enrichment api","email to profile","batch enrichment","\u002Fimg\u002Fblog\u002Fenrich-a-list-from-email-addresses.png","\u002Fimg\u002Fblog\u002Fenrich-a-list-from-email-addresses-card.png",{"path":125,"title":2368,"description":2369,"publishedAt":2338,"author":6,"category":1064,"tags":2370,"readTime":1075,"cover":2373,"listingCover":2374,"image":1069},"A Free API to Extract Page Content for RAG: Read This First","We scraped a Wikipedia page and got 74,552 characters starting with the nav menu. The official API gave 689 clean ones. When each is right.",[1079,1865,2371,2372],"chunking","llm context","\u002Fimg\u002Fblog\u002Ffree-api-extract-page-content-rag.png","\u002Fimg\u002Fblog\u002Ffree-api-extract-page-content-rag-card.png",{"path":2376,"title":2377,"description":2378,"publishedAt":2338,"author":6,"category":1120,"tags":2379,"readTime":1075,"cover":2383,"listingCover":2384,"image":1069},"\u002Fblog\u002Fguides\u002Fgoogle-maps-scraper-alternatives","Google Maps Scraper Alternatives: Building a Local Lead List","Two Google Maps actors, same catalogue, different jobs. One returned nothing on a plausible query. What the fields tell you about a business, measured.",[2380,2381,1615,2382],"google maps scraper","local leads","lead list","\u002Fimg\u002Fblog\u002Fgoogle-maps-scraper-alternatives.png","\u002Fimg\u002Fblog\u002Fgoogle-maps-scraper-alternatives-card.png",{"path":2386,"title":2387,"description":2388,"publishedAt":2338,"author":6,"category":1586,"tags":2389,"readTime":1075,"cover":2394,"listingCover":2395,"image":1069},"\u002Fblog\u002Fguides\u002Fmeta-ad-library-longest-running-ads","Mining the Meta Ad Library with Wireflow: Ninety Days Means It Works","Run duration is the profitability signal competitors publish by accident. How to read it, when it lies, and how to turn a long runner into creative.",[2390,2391,2392,2393],"meta ad library","competitor ads","ad creative","facebook ads","\u002Fimg\u002Fblog\u002Fmeta-ad-library-longest-running-ads.png","\u002Fimg\u002Fblog\u002Fmeta-ad-library-longest-running-ads-card.png",{"path":2397,"title":2398,"description":2399,"publishedAt":2400,"author":6,"category":1092,"tags":2401,"readTime":1075,"cover":2403,"listingCover":2404,"image":1069},"\u002Fblog\u002Fguides\u002Fai-agent-needs-two-integrations","A Model Is Half an Agent: AIHubMix for Models, Monid for Tools","A model gateway gets your agent talking. It still cannot look anything up. How to wire both halves, with AIHubMix on models and Monid on tools.","2026-08-18",[1567,2402,2241,2240],"model gateway","\u002Fimg\u002Fblog\u002Fai-agent-needs-two-integrations.png","\u002Fimg\u002Fblog\u002Fai-agent-needs-two-integrations-card.png",{"path":2406,"title":2407,"description":2408,"publishedAt":2400,"author":6,"category":1092,"tags":2409,"readTime":1075,"cover":2412,"listingCover":2413,"image":1069},"\u002Fblog\u002Fguides\u002Fapi-marketplace-for-ai-agents","The Best API Marketplace for AI Agents in 2026","Marketplaces were built for developers who integrate once. An agent chooses at run time. One ordinary task touched five providers across four steps.",[2410,1567,2240,2411],"api marketplace","tool use","\u002Fimg\u002Fblog\u002Fapi-marketplace-for-ai-agents.png","\u002Fimg\u002Fblog\u002Fapi-marketplace-for-ai-agents-card.png",{"path":2415,"title":2416,"description":2417,"publishedAt":2400,"author":6,"category":1106,"tags":2418,"readTime":1075,"cover":2422,"listingCover":2423,"image":1069},"\u002Fblog\u002Fguides\u002Fbusiness-entity-search-api","Business Entity Search by API: What Exists and What Does Not","State registries are the most searched company lookup and the least available by API. We ran the US endpoint on a private company and it returned nothing.",[2419,2420,2421,1424],"business entity search","company registry","sec edgar","\u002Fimg\u002Fblog\u002Fbusiness-entity-search-api.png","\u002Fimg\u002Fblog\u002Fbusiness-entity-search-api-card.png",{"path":2425,"title":2426,"description":2427,"publishedAt":2400,"author":6,"category":1064,"tags":2428,"readTime":1075,"cover":2432,"listingCover":2433,"image":1069},"\u002Fblog\u002Fguides\u002Fscraper-blocked-what-gets-through","Your Scraper Is Blocked: What Actually Gets Through in 2026","A raw request to a Cloudflare-protected page returns 403. We ran the same URL through a managed endpoint and read what came back, field by field.",[1827,2429,2430,2431],"cloudflare","proxies","blocked","\u002Fimg\u002Fblog\u002Fscraper-blocked-what-gets-through.png","\u002Fimg\u002Fblog\u002Fscraper-blocked-what-gets-through-card.png",{"path":2435,"title":2436,"description":2437,"publishedAt":2400,"author":6,"category":1064,"tags":2438,"readTime":1075,"cover":2440,"listingCover":2441,"image":1069},"\u002Fblog\u002Fguides\u002Fweb-scraping-api-for-ai-agents","The Best Web Scraping API for AI Agents in 2026","An agent cannot pick a scraper from a list it has never seen. What changes when the catalogue is discoverable at run time, measured across three phrasings.",[2439,1567,2240,2258],"web scraping api","\u002Fimg\u002Fblog\u002Fweb-scraping-api-for-ai-agents.png","\u002Fimg\u002Fblog\u002Fweb-scraping-api-for-ai-agents-card.png",{"path":2443,"title":2444,"description":2445,"publishedAt":2446,"author":6,"category":1106,"tags":2447,"readTime":1177,"cover":2453,"listingCover":2454,"image":1069},"\u002Fblog\u002Fguides\u002Fapi-to-find-a-company-website","Best API to Search a Company's Homepage From Its Name","Searching a company's homepage by API is resolution, not search. Which endpoint does it, why fuzzy matches are a feature, and how to pick the right row.","2026-08-17",[2448,2449,2450,2451,2452],"best api search company's homepage","company homepage search api","company website api","company lookup","domain resolution","\u002Fimg\u002Fblog\u002Fapi-to-find-a-company-website.png","\u002Fimg\u002Fblog\u002Fapi-to-find-a-company-website-card.png",{"path":2456,"title":2457,"description":2458,"publishedAt":2446,"author":6,"category":1064,"tags":2459,"readTime":1075,"cover":2463,"listingCover":2464,"image":1069},"\u002Fblog\u002Fguides\u002Fcompany-news-api-press-releases","PR Newswire API: How to Read Releases, Not Send Them","The newswires sell distribution, not access. Reading a company's releases and its coverage is a different endpoint, and one field separates the two.",[2460,2461,2140,2462],"company news api","press release api","signals","\u002Fimg\u002Fblog\u002Fcompany-news-api-press-releases.png","\u002Fimg\u002Fblog\u002Fcompany-news-api-press-releases-card.png",{"path":2466,"title":2467,"description":2468,"publishedAt":2446,"author":6,"category":1106,"tags":2469,"readTime":1177,"cover":2473,"listingCover":2474,"image":1069},"\u002Fblog\u002Fguides\u002Fcrunchbase-api-alternatives-funding-data","Startup Funding Data by API: Crunchbase Alternatives That Bill Per Call","Startup funding data without a Crunchbase seat: one enrichment call carried 33 rounds and 14 stage types on one company, and one field that misleads.",[2470,2471,1421,1660,2472,2106],"startup funding","startup funding data","funding data","\u002Fimg\u002Fblog\u002Fcrunchbase-api-alternatives-funding-data.png","\u002Fimg\u002Fblog\u002Fcrunchbase-api-alternatives-funding-data-card.png",{"path":2476,"title":2477,"description":2478,"publishedAt":2446,"author":6,"category":1106,"tags":2479,"readTime":1177,"cover":2484,"listingCover":2485,"image":1069},"\u002Fblog\u002Fguides\u002Fjob-postings-api-hiring-signals","Jooble API and Job Postings Data: Hiring as a Buying Signal","Jooble's API is for partners, not developers. What a job postings call returns instead, and two fields that turned out to be platform defaults, not data.",[2480,2481,1955,1956,2482,2483],"jooble api","jooble","jobs data","gtm","\u002Fimg\u002Fblog\u002Fjob-postings-api-hiring-signals.png","\u002Fimg\u002Fblog\u002Fjob-postings-api-hiring-signals-card.png",{"path":2487,"title":2488,"description":2489,"publishedAt":2446,"author":6,"category":1106,"tags":2490,"readTime":1075,"cover":2494,"listingCover":2495,"image":1069},"\u002Fblog\u002Fguides\u002Ftechnographic-data-platforms-vs-per-call","Technographic Data Platforms vs One API Call: What You Give Up","We ran a stack detection on a major site and half the fields came back empty. What per-call detection sees, and when a platform earns its price.",[2491,2492,2493,2483],"technographic data","tech stack detection","b2b targeting","\u002Fimg\u002Fblog\u002Ftechnographic-data-platforms-vs-per-call.png","\u002Fimg\u002Fblog\u002Ftechnographic-data-platforms-vs-per-call-card.png",{"path":2497,"title":2498,"description":2499,"publishedAt":2500,"author":6,"category":1184,"tags":2501,"readTime":1075,"cover":2504,"listingCover":2505,"image":1069},"\u002Fblog\u002Fguides\u002Fbest-social-media-scraping-api-2026","What Is the Best API for Social Media Scraping in 2026?","No single best one, because the platforms are not one problem. Which endpoint covers which network, and where a per-call bill differs from per-result.","2026-08-14",[2502,1521,2503,2227],"social media scraping api","tiktok","\u002Fimg\u002Fblog\u002Fbest-social-media-scraping-api-2026.png","\u002Fimg\u002Fblog\u002Fbest-social-media-scraping-api-2026-card.png",{"path":2507,"title":2508,"description":2509,"publishedAt":2500,"author":6,"category":1064,"tags":2510,"readTime":1177,"cover":2514,"listingCover":2515,"image":1069},"\u002Fblog\u002Fguides\u002Fmcp-server-live-web-data-agents","Best MCP Servers: Two Engines, Two Lists That Barely Overlap","Of the eleven servers one engine called best, none fetches live web data. Change the wording and the list changes completely. Both answers, measured.",[2511,2512,1484,1567,2513],"best mcp servers","mcp server list","model context protocol","\u002Fimg\u002Fblog\u002Fmcp-server-live-web-data-agents.png","\u002Fimg\u002Fblog\u002Fmcp-server-live-web-data-agents-card.png",{"path":2517,"title":2518,"description":2519,"publishedAt":2500,"author":6,"category":1092,"tags":2520,"readTime":1099,"cover":2523,"listingCover":2524,"image":1069},"\u002Fblog\u002Fguides\u002Fpay-per-call-data-api-vs-subscription","Which Data API Lets You Pay Per Call Instead of a Subscription?","Metered beats a subscription when your usage is bursty and loses when it is steady. The shapes, the crossover, and the measurement that decides it.",[2342,2521,2410,2522],"data api pricing","metered billing","\u002Fimg\u002Fblog\u002Fpay-per-call-data-api-vs-subscription.png","\u002Fimg\u002Fblog\u002Fpay-per-call-data-api-vs-subscription-card.png",{"path":2526,"title":2527,"description":2528,"publishedAt":2500,"author":6,"category":1106,"tags":2529,"readTime":1177,"cover":2533,"listingCover":2534,"image":1069},"\u002Fblog\u002Fguides\u002Fpeople-data-labs-apollo-zoominfo-alternatives","ZoomInfo Competitors: People Data Labs, Apollo, and What to Buy","ZoomInfo competitors split into three jobs: search, enrich, contract. What each is genuinely best at, and why the price gap is not what it looks like.",[2530,2531,2108,2532,2104],"zoominfo competitors","zoominfo alternatives","apollo alternatives","\u002Fimg\u002Fblog\u002Fpeople-data-labs-apollo-zoominfo-alternatives.png","\u002Fimg\u002Fblog\u002Fpeople-data-labs-apollo-zoominfo-alternatives-card.png",{"path":2536,"title":2537,"description":2538,"publishedAt":2500,"author":6,"category":1184,"tags":2539,"readTime":1099,"cover":2540,"listingCover":2541,"image":1069},"\u002Fblog\u002Fguides\u002Freddit-scraping-api-alternatives","Is There an Alternative to Apify for Scraping Reddit?","Yes, several. The more useful question is why a Reddit scraper returns posts that miss your keywords, because that is a search problem, not a scraper one.",[1613,1612,1614,2340],"\u002Fimg\u002Fblog\u002Freddit-scraping-api-alternatives.png","\u002Fimg\u002Fblog\u002Freddit-scraping-api-alternatives-card.png",{"path":2543,"title":2544,"description":2545,"publishedAt":2546,"author":6,"category":1106,"tags":2547,"readTime":1075,"cover":2550,"listingCover":2551,"image":1069},"\u002Fblog\u002Fguides\u002Fbest-linkedin-scraper-api-2026","What Is the Best LinkedIn Scraper API in 2026?","There is no single best LinkedIn scraper API. There are three jobs, and the honest answer is which endpoint fits which job, and what each one gets wrong.","2026-08-12",[2548,2294,2295,2549],"linkedin scraper api","data","\u002Fimg\u002Fblog\u002Fbest-linkedin-scraper-api-2026.png","\u002Fimg\u002Fblog\u002Fbest-linkedin-scraper-api-2026-card.png",{"path":2553,"title":2554,"description":2555,"publishedAt":2556,"author":6,"category":1106,"tags":2557,"readTime":1075,"cover":2562,"listingCover":2563,"image":1069},"\u002Fblog\u002Fguides\u002Fapollo-scraper","Apollo Scraper: Export Apollo.io Leads Safely by API in 2026","An Apollo scraper extension can get your account flagged. How to pull the same Apollo.io people and company data by metered API, with search free.","2026-08-10",[2558,2559,2560,2561],"apollo scraper","apollo.io","lead-generation","b2b-data","\u002Fimg\u002Fblog\u002Fapollo-scraper.png","\u002Fimg\u002Fblog\u002Fapollo-scraper-card.png",{"path":2565,"title":2566,"description":2567,"publishedAt":2556,"author":6,"category":1184,"tags":2568,"readTime":1177,"cover":2571,"listingCover":2572,"image":1069},"\u002Fblog\u002Fguides\u002Freddit-scraper","Reddit API in 2026: Licensed Access or a Public-Page Scraper","Reddit closed self-service API access in 2025. The honest split between the licensed Data API and a public-page scraper you run metered per call.",[1613,1612,2569,2570,2228],"reddit data api","web-scraping","\u002Fimg\u002Fblog\u002Freddit-scraper.png","\u002Fimg\u002Fblog\u002Freddit-scraper-card.png",{"path":2574,"title":2575,"description":2576,"publishedAt":2556,"author":6,"category":1106,"tags":2577,"readTime":1177,"cover":2584,"listingCover":2585,"image":1069},"\u002Fblog\u002Fguides\u002Ftechnographic-data-api","Technographics by API: Three Sources Said 271, 18 and 0","One domain, one afternoon, three technographic routes. A live detector found nothing, a curated database found eighteen, a sales database found 271.",[2578,2579,2580,2581,2582,2583],"technographics","technographics tool","technographic data api","sales tech stack","tech stack","company-enrichment","\u002Fimg\u002Fblog\u002Ftechnographic-data-api.png","\u002Fimg\u002Fblog\u002Ftechnographic-data-api-card.png",{"path":2587,"title":2588,"description":2589,"publishedAt":2590,"author":6,"category":1184,"tags":2591,"readTime":1075,"cover":2593,"listingCover":2594,"image":1069},"\u002Fblog\u002Fguides\u002Fapify-instagram-scraper","The Apify Instagram Scraper: Which Actor to Use in 2026","Apify has five Instagram scrapers, not one. Which actor fits which job, how to run them without getting blocked, and what two measured runs actually cost.","2026-08-04",[2592,2227,2570,2217],"apify instagram scraper","\u002Fimg\u002Fblog\u002Fapify-instagram-scraper.png","\u002Fimg\u002Fblog\u002Fapify-instagram-scraper-card.png",{"path":2596,"title":2597,"description":2598,"publishedAt":2599,"author":6,"category":1106,"tags":2600,"readTime":1177,"cover":2602,"listingCover":2603,"image":1069},"\u002Fblog\u002Fguides\u002Fbest-linkedin-outreach-tools-2026","The Best LinkedIn Outreach Tools in 2026","LinkedIn outreach tools split into two layers in 2026: the data layer that finds and verifies people, and the sending layer that sequences the messages.","2026-07-29",[2601,2294,2295,2549],"linkedin outreach","\u002Fimg\u002Fblog\u002Fbest-linkedin-outreach-tools-2026.png","\u002Fimg\u002Fblog\u002Fbest-linkedin-outreach-tools-2026-card.png",{"path":2605,"title":2606,"description":2607,"publishedAt":2608,"author":6,"category":1308,"tags":2609,"readTime":2613,"cover":2614,"listingCover":2615,"image":1069},"\u002Fblog\u002Fguides\u002Famazon-pa-api-alternatives","Amazon's PA-API Retires in 2026: How to Move to Monid","PA-API retired in May 2026. Which endpoints actually replace GetItems and SearchItems, tested on the day of writing, including one returning empty prices.","2026-07-21",[2610,2611,2549,2612],"amazon product advertising api","pa-api alternative","ecommerce","9 min","\u002Fimg\u002Fblog\u002Famazon-pa-api-alternatives.png","\u002Fimg\u002Fblog\u002Famazon-pa-api-alternatives-card.png",{"path":2617,"title":2618,"description":2619,"publishedAt":2620,"author":6,"category":1308,"tags":2621,"readTime":1913,"cover":2623,"listingCover":2624,"image":1069},"\u002Fblog\u002Fguides\u002Fbest-amazon-reviews-api-2026","The Best Amazon Reviews API in 2026 (We Tested Them)","There is no official Amazon reviews API. Here are the real ways to get review text, what all of them costs, and the pagination limit nobody mentions.","2026-07-09",[2622,2549,2217,2612],"amazon reviews api","\u002Fimg\u002Fblog\u002Fbest-amazon-reviews-api-2026.png","\u002Fimg\u002Fblog\u002Fbest-amazon-reviews-api-2026-card.png",{"path":2626,"title":2627,"description":2628,"publishedAt":2620,"author":6,"category":1106,"tags":2629,"readTime":1177,"cover":2632,"listingCover":2633,"image":1069},"\u002Fblog\u002Fguides\u002Fbest-email-verification-api-2026","Best Email Verification API in 2026: ZeroBounce, NeverBounce, Kickbox, or One Call on Monid?","ZeroBounce, NeverBounce, Kickbox, Bouncer, a DIY SMTP check and Strale compared: credit packs versus one metered per-call endpoint.",[2630,2549,2217,2631],"email verification api","email","\u002Fimg\u002Fblog\u002Fbest-email-verification-api-2026.png","\u002Fimg\u002Fblog\u002Fbest-email-verification-api-2026-card.png",{"path":2635,"title":2636,"description":21,"publishedAt":2637,"author":1069,"category":1092,"tags":2638,"readTime":1069,"cover":2641,"listingCover":1069,"image":1069},"\u002Fblog\u002Fakta-pro-is-now-available-on-monid","Akta Pro Is Now Available On Monid","2026-07-07",[2217,2639,2640,2549],"partner-tools","private-markets","\u002Fimg\u002Fblog\u002Fakta-pro-is-now-available-on-monid-v2.png",{"path":2643,"title":2644,"description":2645,"publishedAt":2646,"author":1069,"category":1092,"tags":2647,"readTime":1069,"cover":2650,"listingCover":1069,"image":1069},"\u002Fblog\u002Fyour-claude-code-can-now-make-phone-calls","Your Claude Code can now make phone calls","Saperly is now available on Monid. Your agent can now make phone calls for you.","2026-07-05",[2217,2639,2648,2649],"voice","phone","\u002Fimg\u002Fblog\u002Fyour-claude-code-can-now-make-phone-calls.png",{"path":2652,"title":2653,"description":2654,"publishedAt":2655,"author":1069,"category":1092,"tags":2656,"readTime":1069,"cover":2659,"listingCover":1069,"image":1069},"\u002Fblog\u002Fintroducing-suzanne-chatgpt-for-3d-models","Introducing\nClaude for 3D models","Suzanne is now available on Monid. Turn any idea into a production-ready 3D model in one prompt.","2026-06-25",[2657,2217,2658],"3d","creative-tools","\u002Fimg\u002Fblog\u002Fintroducing-suzanne-chatgpt-for-3d-models.png",{"path":2661,"title":2662,"description":2663,"publishedAt":2664,"author":1069,"category":1092,"tags":2665,"readTime":1069,"cover":2668,"listingCover":1069,"image":1069},"\u002Fblog\u002Fminimax-is-now-available-on-monid","MiniMax is now available on Monid","Create images and music with MiniMax models through Monid.","2026-06-24",[2217,2658,2666,2667],"image-generation","music-generation","\u002Fimg\u002Fblog\u002Fminimax-is-now-available-on-monid.png",1791431413892]