DMF_Crawler/docs/research/_raw/02-github-similar.raw.md
Yun Chan 56a6e2da93 chore: 저장소 구조 정리 및 문서화, 첫 커밋
- src/dist 산출물 분리 원칙 정리(.gitignore, .gitattributes)
- 루트 및 주요 폴더(config/scripts/prompts/tests/src, 런타임 폴더 5종)에
  안내용 README.md 추가
- CHANGELOG.md, LICENSE, docs/ops/05-release-and-versioning.md 추가
- docs/README.md 문서 지도 갱신
2026-09-04 09:25:44 +09:00

214 KiB
Raw Permalink Blame History

RAW RESEARCH DUMP — agent-a6dab9e1227de3b58

ORIGINAL TASK PROMPT

오늘 날짜는 2026-09-02 이다. 너는 리서치 에이전트다. 반드시 먼저 ToolSearch 로 "select:WebSearch,WebFetch" 를 로드하고, WebSearch 로 최소 8회 이상 다양한 한국어/영어 질의를 던지고, 핵심 출처 페이지는 WebFetch 로 실제 열어 내용을 확인하라. 실제로 열어 확인한 항목만 verified_by_fetch=true 로 표시하라. 존재를 확인하지 못한 URL, GitHub 저장소, 논문, CLI 플래그는 절대 지어내지 말고 confidence='low' 로 표시하거나 제외하라. 한국 사이트(nedrug.mfds.go.kr, data.go.kr 등)는 WebFetch 가 실패할 수 있으니 실패하면 그 사실을 open_questions 에 적어라. 결과의 summary/detail/recommendations 는 한국어로 쓰되 고유명사·코드·플래그는 원문 유지. 코드 스니펫은 실제 동작 가능한 수준으로 구체적으로 작성하라. 최종 출력은 StructuredOutput 스키마에 맞춰라.

프로젝트 맥락: Windows 11 PC 에서 매일 06:00 에 한국 식약처 원료의약품 등록(DMF) 공고/현황을 크롤링하여 신규/변경/취하 건을 탐지하고, 탭(시트)별로 연동된 보기 좋은 xlsx 리포트를 생성한다. 크롤링·요약 일부를 AI 에이전트 CLI(Claude Code 의 'claude -p' headless 모드 등)로 non-interactive 하게 돌리고, 재부팅 후에도 자동 복구되는 서비스/스케줄러로 운영하며, 서비스가 죽으면 Windows 알림으로 복구 안내를 띄운다.

[축 2: 유사 프로젝트 GitHub 사례] 조사 항목: 다음 범주에서 실제 존재하는 GitHub 저장소를 찾아 URL·스타 수·최근 커밋·언어·핵심 구조·재사용 포인트를 정리하라. 반드시 WebFetch 로 저장소 페이지를 열어 실존을 확인하라. a) 식약처/의약품안전나라(nedrug) 또는 공공데이터포털 의약품 API 크롤러·래퍼 (한국어 저장소 포함: "nedrug crawler", "식약처 크롤링", "의약품 크롤러", "mfds api python") b) FDA DMF / Drugs@FDA / Orange Book / openFDA 스크레이퍼·모니터 c) 규제기관 공고 변경 감지(regulatory change monitoring, pharma regulatory intelligence) 도구 d) 범용 웹 변경 감지 도구: changedetection.io, urlwatch, huginn 등 — 스케줄·diff·알림 구조 e) '매일 크롤링 → 엑셀/구글시트 리포트 → 알림' 파이프라인 예시 (Scrapy/Playwright/requests + openpyxl/xlsxwriter + Task Scheduler/cron) f) Claude Code / Gemini CLI / Codex 를 headless 로 크론 파이프라인에 넣은 예시 저장소 g) Windows 서비스화(NSSM/WinSW/pywin32) + 워치독 + 토스트 알림 예시 h) awesome 리스트(awesome-web-scraping, awesome-pharma-data 등) 각 저장소마다 '이 프로젝트에서 무엇을 가져다 쓸 수 있는가'를 구체적으로 적어라.

dimension 필드에는 "github-similar" 를 넣어라.

[SEARCH #1] github nedrug 식약처 크롤링 의약품 crawler python

Web search results for query: "github nedrug 식약처 크롤링 의약품 crawler python"

Links: [{"title":"GitHub - diakes/coupang_crawler_python · GitHub","url":"https://github.com/diakes/coupang_crawler_python"},{"title":"https://nedrug.mfds.go.kr/index","url":"https://nedrug.mfds.go.kr/index"},{"title":"GitHub - Q00/data.go.kr-crawling: 건강정보, 의약품 크롤링, 멀티쓰레딩 gevent · GitHub","url":"https://github.com/Q00/data.go.kr-crawling"},{"title":"https:/nedrug.mfds.go.kr - 식품의약품안전처","url":"https://nedrug.mfds.go.kr/"},{"title":"의약품안전나라 > 의약품등 검색 - 식품의약품안전처","url":"https://nedrug.mfds.go.kr/searchDrug"},{"title":"Web crawler","url":"https://en.wikipedia.org/wiki/Web_crawler"},{"title":"의약품안전나라 > 사용자별서비스 > 일반소비자","url":"https://nedrug.mfds.go.kr/pbp/CCBRA01"},{"title":"의약품 공공데이터공개","url":"https://nedrug.mfds.go.kr/cntnts/80"},{"title":"네이버 의약품사전 크롤링","url":"https://velog.io/@xenrose/naverPillCrawling"},{"title":"식품안전나라 크롤링 가이드 · 달이 차오른다","url":"https://samslow.github.io/diary/2019/02/11/sickfoom-crawling-guide/"}]

Based on the search results, here's what I found related to your query about nedrug, 식약처 (MFDS), medication crawling, and Python:

Key Resources Found:

  1. NEDRUG Official Portal: NEDRUG is a medication search service operated by the MFDS that includes medication information search, cyber complaints, product and manufacturer information, medication advertisement search, and specification standards.

  2. Public Data API: The MFDS provides high-quality medication information in file data formats (CSV, EXCEL) and OpenAPI form to enable convenient public services.

  3. Related GitHub Projects: The search found references to medication crawling projects on GitHub, including:

    • A data.go.kr crawling project focused on health information and medication crawling with multithreading using gevent
  4. Alternative Approach: There are also examples of crawling pharmacy databases like Naver's medication dictionary using Python libraries such as Selenium and BeautifulSoup.

Recommendation: Rather than web scraping NEDRUG directly, it would be better to use the official public data API that the MFDS provides, as this is the officially supported method for accessing medication information.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #2] github 공공데이터포털 식약처 의약품 API python wrapper mfds

Web search results for query: "github 공공데이터포털 식약처 의약품 API python wrapper mfds"

Links: [{"title":"공공데이터포털","url":"https://www.data.go.kr/dataset/15020626/openapi.do"},{"title":"공공데이터 목록 및 이용안내 - 식의약 데이터 포털","url":"https://data.mfds.go.kr/cntnts/20"},{"title":"k-skill/docs/features/mfds-food-safety.md at main · NomaDamas/k-skill","url":"https://github.com/NomaDamas/k-skill/blob/main/docs/features/mfds-food-safety.md"},{"title":"공공데이터 상세 < 공공데이터 < 외부포털 - 식품의약품안전처","url":"https://data.mfds.go.kr/OPCAA01F01/search?selectedTab=tab1&taskDivsCd=3&taskDivsDtlCd=7&rchSrvcKorNm=&btnSearch="},{"title":"공공데이터 상세 - 식의약 데이터 포털","url":"https://data.mfds.go.kr/OPCAA01F01"},{"title":"식의약 데이터 포털","url":"https://data.mfds.go.kr/"},{"title":"식품의약품안전처 식품의약품안전평가원_식품의약품안전처 연구관리 기술 분류 정보조회 서비스 | 공공데이터포털","url":"https://www.data.go.kr/data/15068423/openapi.do"},{"title":"의약품 공공데이터공개","url":"https://nedrug.mfds.go.kr/cntnts/80"},{"title":"공공데이터 포털API사용하기","url":"https://velog.io/@almondbreez0_3/xiniel0v"}]

Based on the web search results, I found relevant information about MFDS (식품의약품안전처) medication APIs and public data portals:

Key Findings:

Official Data Portals:

  • The MFDS provides pharmaceutical data including medication overview information, medication product approval information, and medical device approval information in XML and JSON formats.
  • Public data is available through the public data portal (data.go.kr) and the pharmaceutical data portal (data.mfds.go.kr).

Available MFDS Medication APIs: Popular MFDS medication APIs include the medication pill identification information API and the medication overview information (e약은요) API.

Python Integration: There is a k-skill project on GitHub that includes MFDS food safety helpers with Python3 scripts for accessing public data APIs, though this example is focused on food safety rather than medications specifically.

The search results show that MFDS medication data is available through official APIs, but I did not find a specific pre-built Python wrapper library for MFDS medication APIs in the search results. You may need to either create your own wrapper or check GitHub repositories for community-maintained wrapper projects by searching for MFDS medication API Python implementations.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #3] github openFDA python client Orange Book Drugs@FDA scraper

Web search results for query: "github openFDA python client Orange Book Drugs@FDA scraper"

Links: [{"title":"FDA Orange Book Scraper | Patent & Exclusivity API API · Apify","url":"https://apify.com/labrat011/fda-orange-book-scraper/api"},{"title":"GitHub - DarpitPatel/OpenFDA: Python code to Scrape data from OpenFDA · GitHub","url":"https://github.com/DarpitPatel/OpenFDA"},{"title":"drugs@fda scraper · GitHub","url":"https://gist.github.com/seanherron/5997278"},{"title":"OpenFDA Drug & Device Data Scraper API in Python · Apify","url":"https://apify.com/fortuitous_pirate/openfda-scraper/api/python"},{"title":"Fda","url":"https://open.fda.gov/apis/drug/orangebook/"},{"title":"openFDA","url":"https://open.fda.gov/"},{"title":"GitHub - coderxio/openfda: Python scripts for capturing OpenFDA data in a database. · GitHub","url":"https://github.com/coderxio/OpenFDA"},{"title":"Approved Drug Products with Therapeutic Equivalence Evaluations","url":"https://en.wikipedia.org/wiki/Approved_Drug_Products_with_Therapeutic_Equivalence_Evaluations"},{"title":"GitHub - FDA/openfda: openFDA is an FDA project to provide open APIs, raw data downloads, documentation and examples, and a developer community for an important collection of FDA public datasets. · GitHub","url":"https://github.com/fda/openfda"},{"title":"Drug API Endpoints","url":"https://open.fda.gov/apis/drug/"}]

Here are the search results for GitHub openFDA Python client Orange Book Drugs@FDA scraper:

Key Resources Found:

  1. DarpitPatel/OpenFDA - A Python project that scrapes data from OpenFDA site using its API and outputs a text file and CSV file with the variables of interest. (GitHub: https://github.com/DarpitPatel/OpenFDA)

  2. coderxio/OpenFDA - Python scripts for capturing OpenFDA data in a database with setup instructions for virtual environments, requirements installation, and Docker compose deployment. (GitHub: https://github.com/coderxio/OpenFDA)

  3. FDA/openfda (Official) - The official repository containing Python pipelines written with Luigi for processing public FDA data sets (drugs, foods, medical devices, and other) into a JSON format that can be loaded into Elasticsearch. (GitHub: https://github.com/fda/openfda)

  4. Apify Scrapers - There are also Apify-based scrapers available:

    • An FDA Orange Book Scraper that searches the FDA's Drugs@FDA database of 28,000+ drug applications — including New Drug Applications (NDAs), Abbreviated New Drug Applications (ANDAs/generics), and Biologics License Applications (BLAs).

Official openFDA API:

The openFDA Drug API provides endpoints for Approved Drug Products with Therapeutic Equivalence Evaluations (Orange Book) and Drugs@FDA which includes most of the drug products approved since 1939.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #4] github pharma regulatory intelligence monitoring tool open source FDA EMA changes

Web search results for query: "github pharma regulatory intelligence monitoring tool open source FDA EMA changes"

Links: [{"title":"AI and the Future of Regulatory Affairs in the U.S. Pharmaceutical Industry | IntuitionLabs","url":"https://intuitionlabs.ai/articles/ai-future-regulatory-affairs-pharma"},{"title":"Open Source Pharma: Tools & Trends in Drug Development | IntuitionLabs","url":"https://intuitionlabs.ai/articles/open-source-pharma-trends"},{"title":"GitHub - anton-semerenko/pharma-radar: Autonomous LLM agent delivering a daily, source-verified pharma regulatory & competitive intelligence briefing to Telegram. · GitHub","url":"https://github.com/anton-semerenko/pharma-radar"},{"title":"Biopharma Regulatory Compliance Services | Clarivate","url":"https://clarivate.com/life-sciences-healthcare/research-development/regulatory-compliance-intelligence/"},{"title":"pharmaceuticals · GitHub Topics","url":"https://github.com/topics/pharmaceuticals?l=python"},{"title":"fda · GitHub Topics · GitHub","url":"https://github.com/topics/fda"},{"title":"open-fda · GitHub Topics · GitHub","url":"https://github.com/topics/open-fda"},{"title":"Food and Drug Administration · GitHub","url":"https://github.com/FDA"},{"title":"Regulatory Intelligence Database, Software & Tools for Compliance","url":"https://www.vistaar.ai/blog/regulatory-intelligence-database-software-tools-for-compliance/"}]

Based on the search results, here are the key findings for pharma regulatory intelligence monitoring tools with open source and FDA/EMA change tracking:

Notable Open Source Tools

Pharma-Radar - An autonomous LLM agent delivering a daily, source-verified pharma regulatory & competitive intelligence briefing to Telegram. It tracks pharma-relevant signals across regulators including FDA, EMA, WHO, trade press, and regulatory registries. Items are verified by construction with at least 2 independent sources or one official regulatory primary before being included.

OpenFDA - OpenFDA is an FDA project to provide open APIs, raw data downloads, documentation and examples, and a developer community for an important collection of FDA public datasets.

Regulatory Intelligence Approaches

AI has emerged as a game-changer for regulatory intelligence, with many pharmaceutical companies leveraging NLP and data mining to monitor FDA databases, guidelines, and publicly available documents for changes that could impact their products.

Recent Regulatory Developments

In January 2026, the FDA and EMA jointly issued "Guiding Principles for Good AI Practice in Drug Development," signaling regulatory openness to new tooling paradigms.

The search results show an active ecosystem of open source and commercial tools designed to help pharmaceutical organizations monitor FDA and EMA regulatory changes in real-time.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #5] changedetection.io github urlwatch huginn web change detection comparison

Web search results for query: "changedetection.io github urlwatch huginn web change detection comparison"

Links: [{"title":"Best Open-Source Website Change Detection Tools (and When Managed Wins) | PageCrawl.io","url":"https://pagecrawl.io/blog/open-source-website-change-detection-tools"},{"title":"GitHub - mattwolfe/changedetection: changedetection.io - The best and simplest self-hosted free open source website change detection monitoring and notification service. An alternative to Visualping, Watchtower etc. Designed for simplicity - the main goal is to simply monitor which websites had a text change for free. Free Open source web page change detection · GitHub","url":"https://github.com/mattwolfe/changedetection"},{"title":"A review of 'changedetection.io,' a free, self-hostable monitoring tool that automatically checks for website changes and notifies you of them. - GIGAZINE","url":"https://gigazine.net/gsc_news/en/20260517-changedetection-io"},{"title":"ChangeDetection.io API","url":"https://changedetection.io/docs/api_v1/index.html"},{"title":"GitHub - dgtlmoon/changedetection.io: Best and simplest tool for website change detection, web page monitoring, and website change alerts. Perfect for tracking content changes, price drops, restock alerts, and website defacement monitoring—all for free or enjoy our SaaS plan!","url":"https://github.com/dgtlmoon/changedetection.io"},{"title":"urlwatch Alternatives","url":"https://alternativeto.net/software/urlwatch"},{"title":"changedetection.io Alternatives","url":"https://alternativeto.net/software/changedetection-io/?p=2"},{"title":"changedetection.io Alternatives","url":"https://alternativeto.net/software/changedetection-io/?p=3"},{"title":"urlwatch Alternatives","url":"https://alternativeto.net/software/urlwatch/?p=5"}]

Based on the search results, here's a comparison of these web change detection tools:

Tool Categories

The tools fall into two categories: purpose-built monitors (changedetection.io, urlwatch) that do change detection and little else, and general automation platforms (Huginn, n8n) where change detection is one workflow among many.

Individual Tools

changedetection.io changedetection.io is a self-hosted free open source website change detection monitoring and notification service, designed for simplicity with the main goal of monitoring which websites had a text change for free. Features include specifying target ranges using CSS selectors, XPath, JSONPath, and jq, filtering using regular expressions, and monitoring text changes in PDF files.

urlwatch urlwatch is intended to help you watch changes in webpages and get notified via e-mail, terminal or through various third party services of any changes, with notifications including the URL that has changed and a unified diff of what has changed.

Huginn Huginn is a system for building agents that perform automated tasks for you online, where agents can read the web, watch for events, and take actions on your behalf. Huginn is noted as the best alternative to changedetection.io, and is both free and open source.

Capability Considerations

When choosing between these tools, important factors to consider include JavaScript rendering, noise filtering, notification breadth, failure visibility, and maintenance burden.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #6] github FDA "drug master file" DMF list scraper monitor

Web search results for query: "github FDA "drug master file" DMF list scraper monitor"

Links: [{"title":"List of Drug Master Files (DMFs) | FDA","url":"https://www.fda.gov/drugs/drug-master-files-dmfs/list-drug-master-files-dmfs"},{"title":"Drug Master File","url":"https://en.wikipedia.org/wiki/Drug_Master_File"},{"title":"GitHub - jbremz/FDA-Analysis: Scraping and analysis of the (now retired) Drugs@FDA site - with scrapy and pandas","url":"https://github.com/jbremz/FDA-Analysis"},{"title":"US Drug Master Files (DMF) Database | PharmaCompass.com","url":"https://www.pharmacompass.com/us-drug-master-files-dmfs"},{"title":"GitHub - logiover/fda-data-scraper: FDA data scraper — openFDA drug/device/food recalls, adverse events & 510(k) clearances as JSON/CSV. No API key, no login.","url":"https://github.com/logiover/fda-data-scraper"},{"title":"GitHub - shaayohn/fda-drug-aproval-data-scraping: This program pulls data from the Drugs@FDA database to give approval details for drugs within a certain criteria · GitHub","url":"https://github.com/shaayohn/fda-drug-aproval-data-scraping"},{"title":"Drug Master Files (DMFs) | FDA","url":"https://www.fda.gov/drugs/forms-submission-requirements/drug-master-files-dmfs"},{"title":"Types of Drug Master Files (DMFs) | FDA","url":"https://www.fda.gov/drugs/drug-master-files-dmfs/types-drug-master-files-dmfs"},{"title":"Guideline for Drug Master Files (DMF) | FDA","url":"https://www.fda.gov/drugs/drug-master-files-dmfs/guideline-drug-master-files-dmf"}]

Based on the search results, here are the relevant findings for GitHub projects related to FDA drug master files and scraping:

GitHub FDA Scraping Projects Found

1. FDA-Analysis (jbremz/FDA-Analysis) This project contains all of the code used to scrape the data from the (now retired) Drugs@FDA site and analyse it. The British Medical Journal asked the creator to build a web scraper to crawl through over 22,000 different products, and it uses Scrapy and pandas for analysis.

2. FDA Data Scraper (logiover/fda-data-scraper) FDA data scraper — openFDA drug/device/food recalls, adverse events & 510(k) clearances as JSON/CSV. This tool is designed to scrape FDA recalls, adverse events, and 510(k) clearances data with the public openFDA API, which is accessible without an API key, account or login.

3. FDA Drug Approval Data Scraping (shaayohn/fda-drug-aproval-data-scraping) This program pulls data from the Drugs@FDA database to give approval details for drugs within a certain criteria.

Official FDA DMF Resources

The list of DMFs, which is updated quarterly, contains DMFs received by June 30, 2026, for which acknowledgment letters were sent before July 19, 2026. The official FDA provides info on FDA's list of Drug Master Files (DMFs) and info on submission of DMFs to agency, including DMF list, status of DMF, types of DMF, and more.

Note: The search results show several GitHub projects for FDA data scraping, but none specifically focused on Drug Master Files (DMF) list scraping/monitoring. The projects found target broader FDA data like recalls, adverse events, and general drug approvals.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #7] github daily scraper openpyxl excel report Windows Task Scheduler python example

Web search results for query: "github daily scraper openpyxl excel report Windows Task Scheduler python example"

Links: [{"title":"scheduled-tasks · GitHub Topics · GitHub","url":"https://github.com/topics/scheduled-tasks?l=python&o=desc&s=updated"},{"title":"GitHub - prabudevarajan/Task-Reminder-Automation-Python-Excel-CSV-Email-Alerts: A Python-based task reminder tool with a simple Tkinter UI. Tasks are stored in Excel/CSV and email reminders are sent automatically at 15, 7, 3, and 1 day(s) before the due date. Includes customizable email templates, daily scheduler, and logging. Free and open source. · GitHub","url":"https://github.com/prabudevarajan/Task-Reminder-Automation-Python-Excel-CSV-Email-Alerts"},{"title":"GitHub - Bwhiz/Auto-Excel-Reports: This repository contains scripts and workflows to automate the generation and distribution of Excel reports using Python's openpyxl library. The automation is achieved using GitHub Actions, enabling scheduled and event-triggered report generation and email distribution. · GitHub","url":"https://github.com/Bwhiz/Auto-Excel-Reports"},{"title":"openpyxl · GitHub Topics · GitHub","url":"https://github.com/topics/openpyxl?o=asc&s=forks"},{"title":"openpyxl-python · GitHub Topics · GitHub","url":"https://github.com/topics/openpyxl-python"},{"title":"GitHub - god233012yamil/Excel-Automation-Using-Python: Repository for automating Excel tasks using Python and the "openpyxl" library · GitHub","url":"https://github.com/god233012yamil/Excel-Automation-Using-Python"},{"title":"excelwriter · GitHub Topics","url":"https://github.com/topics/excelwriter?l=python"},{"title":"python-excel · GitHub Topics · GitHub","url":"https://github.com/topics/python-excel"}]

I found several relevant GitHub repositories related to your search query. Here are the key findings:

Automated Excel Report Generation Projects

Auto-Excel-Reports - This repository contains scripts and workflows to automate the generation and distribution of Excel reports using Python's openpyxl library. The automation is achieved using GitHub Actions, enabling scheduled and event-triggered report generation and email distribution. It uses openpyxl to create and format Excel reports and automatically sends the generated reports via email using SMTP.

Excel Automation Using Python - This repository is designed for automating Excel tasks using Python and the openpyxl library, containing examples and guides to help leverage openpyxl for efficient and effective Excel file manipulation. openpyxl is a powerful library that allows you to read, write, and modify Excel files in Python.

Task Reminder Automation - This project uses libraries including pandas, openpyxl, and schedule. Tasks are stored in Excel/CSV and email reminders are sent automatically at 15, 7, 3, and 1 day(s) before the due date.

Scheduling Approaches

The search results show different scheduling approaches:

  • GitHub Actions workflows can run scripts daily or on push events
  • Scripts could be scheduled as Cron jobs on Linux to automate report population and distribution

While the search results don't specifically showcase Windows Task Scheduler examples, the Python scripts using openpyxl can be integrated with Windows Task Scheduler for automation on Windows systems.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #8] github awesome-web-scraping awesome pharma data list

Web search results for query: "github awesome-web-scraping awesome pharma data list"

Links: [{"title":"GitHub - lorien/awesome-web-scraping: List of libraries, tools and APIs for web scraping and data processing. · GitHub","url":"https://github.com/lorien/awesome-web-scraping"},{"title":"awesome-web-scraping/README.md at master · lorien/awesome-web-scraping","url":"https://github.com/lorien/awesome-web-scraping/blob/master/README.md"},{"title":"GitHub - noirquant/awesome-web-scraping: List of libraries, tools and APIs for web scraping and data processing. · GitHub","url":"https://github.com/noirquant/awesome-web-scraping"},{"title":"GitHub - suriyadeepan/WebScraping-for-Healthcare: Web Scraping for Healthcare · GitHub","url":"https://github.com/suriyadeepan/WebScraping-for-Healthcare"},{"title":"GitHub - jjwangnlp/awesome-web-scraping: List of libraries, tools and APIs for web scraping and data processing.","url":"https://github.com/jjwangnlp/awesome-web-scraping"},{"title":"GitHub - luminati-io/Awesome-Web-Scraping: A list of libraries, tools, and APIs for web scraping and data processing. Find everything you need for extracting, managing, and processing data from the web, from HTTP libraries to browser automation tools and proxy services.","url":"https://github.com/luminati-io/Awesome-Web-Scraping"},{"title":"pharmaceutical-data · GitHub Topics · GitHub","url":"https://github.com/topics/pharmaceutical-data"},{"title":"GitHub - spinov001-art/awesome-web-scraping-2026: 130+ web scraping tools — Python, JavaScript, Go, Rust. Anti-detection, proxies, cloud platforms. Updated weekly. Includes free API alternatives.","url":"https://github.com/spinov001-art/awesome-web-scraping-2026"},{"title":"GitHub - duyet/awesome-web-scraper: A collection of awesome web scaper, crawler. · GitHub","url":"https://github.com/duyet/awesome-web-scraper"},{"title":"GitHub - arpitamangal/pharma-scrape-and-analysis: Improve access to healthcare services and reduce costs. · GitHub","url":"https://github.com/arpitamangal/pharma-scrape-and-analysis"}]

Based on the search results, here's what I found:

Main Results

The awesome-web-scraping repository is a list of packages, services, and manuals related to web scraping. The primary repository is maintained by lorien at lorien/awesome-web-scraping on GitHub.

Pharmaceutical Data Scraping

There's a dedicated project called "WebScraping-for-Healthcare" focused on scraping the internet for extracting healthcare and pharma data. This repository appears to be one of the more relevant results for pharmaceutical data specifically.

Additionally, there are several other relevant repositories:

  • An AI-powered pipeline for scraping and tracking pharmaceutical clinical pipeline data with LLM-based change detection
  • A pharma-scrape-and-analysis project that identifies substitutes from lesser-known brands and classifies medicines into FDA's list of 40 categories

General Web Scraping Resources

There are comprehensive lists of libraries, tools, and APIs for web scraping and data processing, including HTTP libraries, browser automation tools, and proxy services.

The search results show multiple "awesome" list variations on GitHub focused on web scraping tools, with some specifically tailored to pharmaceutical and healthcare data extraction.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #9] github python windows service nssm winsw watchdog toast notification win10toast

Web search results for query: "github python windows service nssm winsw watchdog toast notification win10toast"

Links: [{"title":"GitHub - ysfchn/toasted: Rich toast notifications library for Windows (10/11) in Python, supports all elements provided by Windows. · GitHub","url":"https://github.com/ysfchn/toasted"},{"title":"GitHub - jithurjacob/Windows-10-Toast-Notifications: Python library to display Windows 10 Toast Notifications · GitHub","url":"https://github.com/jithurjacob/Windows-10-Toast-Notifications"},{"title":"GitHub - DatGuy1/Windows-Toasts: Python library used to send toast notifications on Windows machines · GitHub","url":"https://github.com/DatGuy1/Windows-Toasts"},{"title":"Generating toast notifications on windows 10 using Python · GitHub","url":"https://gist.github.com/hygull/32a742339a416dcfa2990504c848c1a9"},{"title":"Windows-10-Toast-Notifications/win10toast at master · jithurjacob/Windows-10-Toast-Notifications","url":"https://github.com/jithurjacob/Windows-10-Toast-Notifications/tree/master/win10toast"},{"title":"GitHub - larsekje/PythonWindowsServices: A proof-of-concept on how to run a python script as a windows service using nssm · GitHub","url":"https://github.com/larsekje/PythonWindowsServices"},{"title":"Windows-10-Toast-Notifications/win10toast/init.py at master · jithurjacob/Windows-10-Toast-Notifications","url":"https://github.com/jithurjacob/Windows-10-Toast-Notifications/blob/master/win10toast/init.py"},{"title":"GitHub - GitHub30/win11toast: Toast notifications for Windows 10 and 11 based on WinRT · GitHub","url":"https://github.com/GitHub30/win11toast"},{"title":"Windows 10 Toast Notifications","url":"https://github.com/jacobcolbert/Windows-10-Toast-Notifications"},{"title":"An easy-to-use Python library for displaying Windows 10 Toast Notifications","url":"https://pypi.org/project/win10toast/"}]

I found several relevant GitHub repositories and resources related to your search query:

Windows Service Management

NSSM (the Non-Sucking Service Manager) allows easy creation of windows services with a few simple commands, such as: nssm install "SERVICE_NAME" "PATH_TO_PYTHON.exe" "PATH_TO_PYTHON_SCRIPT.py". There's a proof-of-concept project (PythonWindowsServices) on GitHub demonstrating how to run a Python script as a Windows service using NSSM.

Toast Notifications Libraries

Several Python libraries are available for Windows 10/11 toast notifications:

  1. win10toast: A Python library to display Windows 10 Toast Notifications, which can be installed via pip and includes features for custom icons and threaded notifications.

  2. toasted: A rich toast notifications library for Windows (10/11) in Python that supports all elements provided by Windows.

  3. win11toast: Toast notifications for Windows 10 and 11 based on WinRT.

  4. Windows-Toasts: A Python library used to send toast notifications on Windows machines.

The search results show these tools work together for creating Python-based Windows services with notification capabilities. The win10toast library appears to be the most commonly used for simple notification implementation.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #10] github "claude -p" headless cron pipeline automation script

Web search results for query: "github "claude -p" headless cron pipeline automation script"

Links: [{"title":"What Is Claude Code Headless Mode? How to Run AI Agents Without a Terminal | MindStudio","url":"https://www.mindstudio.ai/blog/claude-code-headless-mode-autonomous-agents"},{"title":"Claude Code Can Work While You Sleep | wmedia.es","url":"https://wmedia.es/en/tips/claude-code-headless-mode-autonomous-agent"},{"title":"Claude Code in CI/CD and Headless Automation - Running the Agent Unattended in Pipelines | hidekazu-konishi.com","url":"https://hidekazu-konishi.com/entry/claude_code_cicd_and_headless_automation.html"},{"title":"Chapter 17: Headless Mode — Claude as a One-Shot Command · Claude Code for Clinicians","url":"https://iyadsultan.github.io/claude-code-for-clinicians/ch17-headless-mode/"},{"title":"Claude Code in Scripts: Headless Mode and Automation (2026) | StackNotice","url":"https://stacknotice.com/blog/claude-code-headless-scripting-2026"},{"title":"How to Set Up Cron Jobs with Claude Code (Automated Scheduling Guide) - Usagebar Blog","url":"https://usagebar.com/blog/how-to-do-cron-job-setup-on-claude-code"},{"title":"Claude Code Automation: Non-Interactive Mode | DevShelfHub","url":"https://www.devshelfhub.com/tutorials/claude-code/automation/"},{"title":"Claude Code Process Automation — HeyClaude","url":"https://heyclau.de/entry/guides/business-process-automation"},{"title":"Claude Code Headless Mode Guide (2026) | Like One","url":"https://likeone.ai/blog/claude-code-headless-mode-guide-2026/"},{"title":"Claude Code Headless Mode | Build This Now","url":"https://www.buildthisnow.com/blog/guide/development/claude-code-headless-mode"}]

The search results cover Claude Code's headless mode functionality extensively. Here's what I found:

Claude Code Headless Mode with -p Flag

Claude Code headless mode uses the -p flag to accept a prompt directly from the command line and execute without any user interaction. This makes it suitable for automation, scheduled tasks, and multi-agent workflows where human input isn't available — or isn't wanted.

How It Works

In headless mode, you pass a prompt, Claude runs it, prints the result, and exits with no conversation or interactivity — pure input/output. When you pass --print (or -p) with a prompt, Claude Code runs the task, prints output to stdout, and exits with a zero or non-zero exit code depending on success.

Automation and Cron Integration

With the -p flag, Claude Code becomes a headless agent you can schedule with cron. Cron plus -p is how Claude Code goes from "an interactive tool I open when I sit down" to "a service that runs whether I am at your desk or not."

Key Features for Pipelines

The -p flag pairs with --output-format, --max-turns, --model, --allowedTools, and session-continuation flags to make the run controllable and parseable. Headless mode reads stdin and writes stdout, so it composes with any shell pipeline, script, or CI step.

Setup Requirements

Claude Code needs a valid Anthropic API key, and for automation contexts like cron jobs or CI pipelines, this environment variable needs to be available in the execution environment.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #11] github 원료의약품 DMF 등록 식약처 크롤러 파이썬

Web search results for query: "github 원료의약품 DMF 등록 식약처 크롤러 파이썬"

Links: [{"title":"등록대상 원료의약품(DMF) 등록 공고(7월 둘째주) < 법령 및 고시>의약품 인허가정보>제약글로벌정보센터-제약산업정보포털","url":"https://www.khidi.or.kr/board/view?pageNum=48&rowCnt=10&menuId=MENU01872&maxIndex=00487793189998&minIndex=00487441479998&schType=0&schText=&categoryId=&continent=&country=&upDown=0&boardStyle=&no1=912&linkId=26604564"},{"title":"식품의약품안전처_원료의약품등록(DMF)현황 | 공공데이터포털","url":"https://www.data.go.kr/data/15057075/openapi.do"},{"title":"[지침]원료의약품 등록(DMF) 처리 절차 < 법령 및 고시>의약품 인허가정보>제약글로벌정보센터-제약산업정보포털","url":"https://www.khidi.or.kr/board/view?pageNum=1&rowCnt=10&menuId=MENU01872&maxIndex=99999999999999&minIndex=99999999999999&schType=0&schText=&categoryId=&continent=&country=&upDown=0&boardStyle=&no1=0&linkId=26605812"},{"title":"의약품안전나라 > 의약품등 정보 > 의약품 및 화장품 품목정보 > 원료의약품등록(DMF) 정보","url":"https://nedrug.mfds.go.kr/bbs/117"},{"title":"Drug Master File","url":"https://en.wikipedia.org/wiki/Drug_Master_File"},{"title":"사업소개>의약품>허가심사>KDMF | 식품의약품안전평가원","url":"https://www.nifds.go.kr/brd/m_87/list.do"},{"title":"Censorship of GitHub","url":"https://en.wikipedia.org/wiki/Censorship_of_GitHub"},{"title":"DMF","url":"https://en.wikipedia.org/wiki/DMF"},{"title":"GitHub - Q00/data.go.kr-crawling: 건강정보, 의약품 크롤링, 멀티쓰레딩 gevent · GitHub","url":"https://github.com/Q00/data.go.kr-crawling"}]

검색 결과를 찾았습니다. 쿼리와 관련된 몇 가지 결과가 있습니다:

주요 결과:

  1. 공공 API 정보: 식품의약품안전처의 "원료의약품등록(DMF)현황" OpenAPI는 등록된 원료의약품의 현황을 조회할 수 있으며, 등록번호, 성분명, 업체명, 제조소명, 제조소 소재지, 제조국가명, 발급일자 등의 정보를 제공합니다.

  2. GitHub 크롤러 프로젝트: 검색 결과에는 "건강정보, 의약품 크롤링, 멀티쓰레딩 gevent"라는 설명의 GitHub 저장소가 있습니다. 이는 data.go.kr에서 의약품 데이터를 크롤링하는 Python 프로젝트로 보입니다.

  3. 공식 정보: 원료의약품 등록 현황(신규등록, 변경등록, 연차보고) 정보는 식약처 홈페이지 전자민원창구에서 확인할 수 있습니다.

더 구체적인 GitHub 저장소를 찾으시려면, GitHub 사이트에서 직접 검색하거나 위의 Q00/data.go.kr-crawling 저장소를 확인해보시기 바랍니다.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #12] github thp/urlwatch huginn/huginn repository stars

Web search results for query: "github thp/urlwatch huginn/huginn repository stars"

Links: [{"title":"GitHub - thp/urlwatch: Watch (parts of) webpages and get notified when something changes via e-mail, on your phone or via other means. Highly configurable. · GitHub","url":"https://github.com/thp/urlwatch"},{"title":"GitHub - huginn/huginn: Create agents that monitor and act on your behalf. Your agents are standing by! · GitHub","url":"https://github.com/huginn/huginn"},{"title":"simbakeila123 - Overview","url":"https://github.com/simbakeila123"},{"title":"GitHub - huginn/huginn_agent: Base for creating new Huginn Agents as Gems · GitHub","url":"https://github.com/huginn/huginn_agent"},{"title":"Add support for watching GitHub repositories · Issue #246 · thp/urlwatch","url":"https://github.com/thp/urlwatch/issues/246"},{"title":"huginn · GitHub Topics · GitHub","url":"https://github.com/topics/huginn?o=asc&s=stars"},{"title":"github.com","url":"https://github.com/roxwize/huginn"},{"title":"github.com","url":"https://github.com/itkevin/huginn"}]

Based on the search results, I found information about both repositories:

thp/urlwatch: The repository has 3.1k stars and urlwatch is intended to help you watch changes in webpages and get notified (via e-mail, in your terminal or through various third party services) of any changes.

huginn/huginn: The search results don't display the specific star count for this repository in the excerpts provided. However, Huginn is a system for building agents that perform automated tasks for you online. They can read the web, watch for events, and take actions on your behalf.

To get the exact current star counts for both repositories, you would need to visit the GitHub pages directly at https://github.com/thp/urlwatch and https://github.com/huginn/huginn.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #13] github 의약품안전나라 e약은요 API 파이썬 오픈API 크롤링 selenium 의약품 허가

Web search results for query: "github 의약품안전나라 e약은요 API 파이썬 오픈API 크롤링 selenium 의약품 허가"

Links: [{"title":"식의약 데이터 포털","url":"https://data.mfds.go.kr/"},{"title":"식품의약품안전처_의약품개요정보(e약은요) | 공공데이터포털","url":"https://www.data.go.kr/data/15075057/openapi.do"},{"title":"식품의약품안전처 데이터활용서비스","url":"https://www.foodsafetykorea.go.kr/apiMain.do"},{"title":"Censorship of GitHub","url":"https://en.wikipedia.org/wiki/Censorship_of_GitHub"},{"title":"식품의약품안전처_의약품 낱알식별 정보 | 공공데이터포털","url":"https://www.data.go.kr/data/15057639/openapi.do"},{"title":"의약품안전나라 > 공공데이터 정보 > 공공데이터 > 개요","url":"https://nedrug.mfds.go.kr/cntnts/80"},{"title":"공공데이터 포털","url":"https://www.data.go.kr/dataset/15020627/openapi.do"},{"title":"데이터활용서비스","url":"https://www.foodsafetykorea.go.kr/api/openApiAplcInfo.do"},{"title":"GitHub - betagouv/api-medicaments: Base de données publique officielle des médicaments de l'Agence nationale de sécurité du médicament et des produits de santé · GitHub","url":"https://github.com/betagouv/api-medicaments"},{"title":"공공데이터 포털API사용하기","url":"https://velog.io/@almondbreez0_3/xiniel0v"}]

검색 결과를 정리하면 다음과 같습니다:

의약품 데이터 관련 공개 API

식품의약품안전처의 「의약품개요정보(e약은요)」 OpenAPI는 일반의약품 중 공급실적이 있는 제품들의 주요 정보를 제공하는 공공데이터입니다. 주요 제공 항목에는 업체명, 제품명, 품목기준코드, 효능, 사용법, 주의사항, 상호작용, 부작용, 보관법 등이 포함되어 있습니다.

데이터 제공 형식

식품의약품안전처가 보유·관리하고 있는 의약품 정보를 파일 데이터(CSV, EXCEL) 및 OpenAPI 형태로 제공합니다.

관련 API 목록

검색 결과에서 다음과 같은 공공데이터포털 API들을 활용할 수 있습니다:

Python 활용 예시

공공데이터 포털의 Open API를 적용하여 의약품 데이터를 활용하는 프로젝트 사례가 있습니다. 이러한 API들은 requests 라이브러리를 통해 Python에서 직접 호출할 수 있으며, 발급받은 API 키를 통해 JSON 형식의 데이터를 받을 수 있습니다.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #14] github repository Claude Code scheduled cron job "claude -p" daily report automation example repo

Web search results for query: "github repository Claude Code scheduled cron job "claude -p" daily report automation example repo"

Links: [{"title":"Claude Code GitHub Actions - Claude Code Docs","url":"https://code.claude.com/docs/en/github-actions"},{"title":"GitHub - jshchnz/claude-code-scheduler: Put Claude on autopilot · GitHub","url":"https://github.com/jshchnz/claude-code-scheduler"},{"title":"Claude Code in GitHub Actions: Automate Standard Tasks","url":"https://jannikreinhard.com/claude-code-github-actions/"},{"title":"Claude Code Routines — The Cron Replacement I Didnt Know I Needed | by Mayank Jain | Level Up Coding","url":"https://levelup.gitconnected.com/claude-code-routines-the-cron-replacement-i-didnt-know-i-needed-6f53cf476577?gi=f0e7b272cd2b"},{"title":"Complete Guide to Claude Code Scheduled Execution - Automation with GitHub Actions Scheduled Workflows - SmartScope","url":"https://smartscope.blog/en/ai-development/claude-code-scheduled-automation-guide/"},{"title":"🕘 Claude Code + Cron Automation Complete Guide 2025 - Scheduled Execution & Scheduling Implementation - SmartScope","url":"https://smartscope.blog/en/generative-ai/claude/claude-code-cron-schedule-automation-complete-guide-2025/"},{"title":"Complete Guide to Claude Code Scheduled Execution - Automate with GitHub Actions Scheduled Workflows - SmartScope","url":"https://smartscope.blog/en/generative-ai/claude/claude-code-scheduled-automation-guide/"},{"title":" Claude Code × Cron Complete Automation Guide - Streamline Tasks with Scheduled Execution - SmartScope","url":"https://smartscope.blog/en/generative-ai/claude/claude-code-cron-automation-guide/"},{"title":"How to Run Claude Code as a Cron Job - Drew Bredvick","url":"https://drew.tech/posts/claude-code-as-a-cron-job"}]

Based on the search results, I found several resources related to Claude Code scheduled cron job automation for daily report generation:

Key Resources Found:

  1. Official Claude Code Documentation: Claude Code's GitHub Action runs in automation mode on any GitHub event, including a cron schedule. The official docs provide an example workflow that generates a report in the workflow run log at 09:00 UTC each day.

  2. Example GitHub Repository - claude-code-scheduler: The repository jshchnz/claude-code-scheduler demonstrates automated scheduling. At the scheduled time, it runs claude -p "your command" with configuration for daily code reviews and automated tasks.

  3. Claude Code Routines: Routines can be scheduled to run on a recurring cadence - hourly, daily, weekly, or a custom cron expression, offering cloud-based automation for Claude Code.

  4. Practical Implementation Examples: Multiple guides demonstrate using GitHub Actions with cron schedules. One example uses a cron trigger where Claude summarizes repository activity every morning, posting the summary as a new issue titled 'Daily report'.

  5. Local Cron Setup: There are also examples of setting up Claude Code with system cron jobs for daily code review at 9:00 AM or generating reports on specific schedules.

The search results show multiple approaches ranging from GitHub Actions workflows to Claude Code Routines to local system cron jobs, all supporting automated daily report generation using Claude Code.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #15] github WinSW winsw/winsw python script service example xml

Web search results for query: "github WinSW winsw/winsw python script service example xml"

Links: [{"title":"WinSW-Windows (WinSW) · GitHub","url":"https://github.com/WinSW-Windows"},{"title":"winsw/samples/minimal.xml at v3 · winsw/winsw","url":"https://github.com/winsw/winsw/blob/v3/samples/minimal.xml"},{"title":"GitHub - winsw/winsw: A wrapper executable that can run any executable as a Windows service, in a permissive license. · GitHub","url":"https://github.com/winsw/winsw"},{"title":"winsw/samples/complete.xml at v3 · winsw/winsw","url":"https://github.com/winsw/winsw/blob/v3/samples/complete.xml"},{"title":"winsw/docs/xml-config-file.md at v3 · winsw/winsw","url":"https://github.com/winsw/winsw/blob/v3/docs/xml-config-file.md"},{"title":"How to Create a Windows Service with WinSW: A Step-by-Step Guide","url":"https://blog.usro.net/2024/10/how-to-create-a-windows-service-with-winsw-a-step-by-step-guide/"},{"title":"PowerShell Guide: Script as a Windows Service | all posts","url":"https://www.ehmiiz.se/blog/ps_scriptasaservice/"},{"title":"Run any program as a Windows background service with WinSW - Infonautics Blog","url":"https://www.infonautics.ch/blog/run-any-program-as-a-windows-background-service-with-winsw/"}]

I've found relevant search results for WinSW and Python script service examples with XML configuration. Here's what I found:

WinSW Overview

WinSW is a wrapper executable that can run any executable as a Windows service, in a permissive license, and wraps and manages any application as a Windows service.

Basic Setup Steps

To set up a service, take WinSW.exe or WinSW.zip from the distribution, write a myapp.xml configuration file, run winsw install myapp.xml to install the service, run winsw start myapp.xml to start the service, and run winsw status myapp.xml to see if your service is up and running.

XML Configuration

Alternatively, you can rename the WinSW.exe to your taste (such as myapp.exe), write myapp.xml, and place those two files side by side, because that's how WinSW discovers its configuration.

Sample XML Files

The search results reference example configuration files available on GitHub:

  • A minimal Windows Service Wrapper configuration example with only mandatory options
  • A complete example with all configuration options

Configuration XML files can include environment variable expansions of the form %Name%, which will be automatically replaced by the actual values of the variables.

For detailed Python-specific examples, you may want to check the GitHub repository directly at https://github.com/winsw/winsw and review the sample XML files in their /samples directory.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #16] github pywin32 win32serviceutil python windows service example auto restart watchdog

Web search results for query: "github pywin32 win32serviceutil python windows service example auto restart watchdog"

Links: [{"title":"[python-win32] automatically restart python service after crash","url":"https://mail.python.org/pipermail/python-win32/2017-January/013807.html"},{"title":"pywin32/win32/Lib/win32serviceutil.py at main · mhammond/pywin32","url":"https://github.com/mhammond/pywin32/blob/main/win32/Lib/win32serviceutil.py"},{"title":"An example Windows service implemented with pywin32 wrappers. #python #windows-service #pywin32 · GitHub","url":"https://gist.github.com/drmalex07/10554232"},{"title":"Pywin32/lib/x32/win32/lib/win32serviceutil.py at master · SublimeText/Pywin32","url":"https://github.com/SublimeText/Pywin32/blob/master/lib/x32/win32/lib/win32serviceutil.py"},{"title":"[Feature Request] Option to terminate Service and trigger recovery actions · Issue #1563 · mhammond/pywin32","url":"https://github.com/mhammond/pywin32/issues/1563"},{"title":"Python-Windows-Service-Example/example_service.py at master · HaroldMills/Python-Windows-Service-Example","url":"https://github.com/HaroldMills/Python-Windows-Service-Example/blob/master/example_service.py"},{"title":"pywin32/win32/Demos/service/serviceEvents.py at main · mhammond/pywin32","url":"https://github.com/mhammond/pywin32/blob/main/win32/Demos/service/serviceEvents.py"},{"title":"How to Monitor Windows Services Using Python on Windows - Woteq Zone","url":"https://woteq.com/how-to-monitor-windows-services-using-python-on-windows"},{"title":"DEV Community","url":"https://dev.to/demola12/building-a-robust-windows-service-in-python-with-win32serviceutil-part-13-1k6k"}]

I've completed the web search for your query. Here are the key findings:

Service Auto-Restart and Configuration: For services that crash, you can configure the service to restart using ChangeServiceConfig2. The configuration uses a dictionary with failure actions that can specify multiple restart attempts with delays.

Service Watchdog/Monitoring: You can write a monitor script that watches a service and attempts to restart it if it's found stopped, running indefinitely until interrupted. The win32service module allows you to connect to the Service Control Manager, enumerate services, check their state (Running, Stopped, Paused), and start, stop, pause, or restart services.

Basic Service Structure: The script checks if it's being run as a standalone application, and if so, initializes and starts the service using servicemanager. The results include examples from:

  • The main pywin32 repository on GitHub
  • A gist with a complete pywin32 service example
  • A Python Windows Service example repository with implementation details

These resources provide practical examples of implementing Windows services in Python with auto-restart capabilities and service monitoring/watchdog functionality.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #17] github playwright python scraper xlsxwriter daily report telegram slack notification pipeline

Web search results for query: "github playwright python scraper xlsxwriter daily report telegram slack notification pipeline"

Links: [{"title":"playwright-python · GitHub Topics · GitHub","url":"https://github.com/topics/playwright-python?o=asc&s=updated"},{"title":"Reporter to send the result of Playwright Automation to Slack | Medium","url":"https://medium.com/@indraaristya/playwright-report-to-slack-e07e8996c9de"},{"title":"GitHub - ManiMozaffar/linkedIn-scraper: A playwright bot which is implemented to scrape linkedin and store advertisement data in a database and telegram channel · GitHub","url":"https://github.com/ManiMozaffar/linkedIn-scraper"},{"title":"GitHub - dineshk-qa/playwright.slack.reporter: slack reporter for playwright framework","url":"https://github.com/dineshk-qa/playwright.slack.reporter"},{"title":"xlsxwriter · GitHub Topics · GitHub","url":"https://github.com/topics/xlsxwriter?l=python"},{"title":"Playwright with GitHub Actions and Slack Notification | by Vinayak M | Medium","url":"https://medium.com/@vinayakhk9/playwright-with-github-actions-and-slack-notification-b56eb982659b"},{"title":"GitHub - HasData/playwright-scraping: This repository demonstrates web scraping and browser automation using Playwright in both Python and Node.js. It includes scripts for common tasks such as scraping data, interacting with web elements, handling authentication, and managing errors. · GitHub","url":"https://github.com/HasData/playwright-scraping"},{"title":"playwright · GitHub Topics · GitHub","url":"https://github.com/topics/playwright?l=python"}]

I found several relevant GitHub projects and resources related to your search query. Here's what the results show:

Web Scraping with Playwright & Python

There are repositories that demonstrate web scraping and browser automation using Playwright in both Python and Node.js, including scripts for common tasks such as scraping data, interacting with web elements, handling authentication, and managing errors.

Slack & Telegram Notification Integration

There are custom Playwright reporters that allow sending automation results to Slack channels, which count how many tests passed, failed, and were flaky, then send it directly to the Slack webhook.

For Telegram integration, Playwright-based scrapers can send collected data to a Telegram chat or channel, making it easy to share with your community or track for personal use.

LinkedIn Scraper Example

One notable project uses Playwright for advanced browser automation and the FastAPI framework to build a lightweight, high-performance API for retrieving information quickly and efficiently.

Report Generation & Excel Export

Projects exist that scrape data and save it into Excel files, with tools like xlsxwriter and pandas commonly used for data frame management and spreadsheet generation.

The search results include various GitHub repositories with implementations combining these technologies for automated reporting pipelines with notification capabilities.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #18] github pharmaceutical pipeline scraping LLM change detection tracker

Web search results for query: "github pharmaceutical pipeline scraping LLM change detection tracker"

Links: [{"title":"pharmaceutical-data · GitHub Topics · GitHub","url":"https://github.com/topics/pharmaceutical-data"},{"title":"GitHub - patrickloeber/llm-data-scrapers: A list of useful Open Source tools and scrapers to gather data for LLMs · GitHub","url":"https://github.com/patrickloeber/llm-data-scrapers"},{"title":"GitHub - FareedKhan-dev/best-llm-finder-pipeline: Agentic RAG, Multi-Agent Systems, and Vision Reasoning are three pipelines to find the perfect LLM · GitHub","url":"https://github.com/FareedKhan-dev/best-llm-finder-pipeline"},{"title":"GitHub - architkaila/Fine-Tuning-LLMs-for-Medical-Entity-Extraction: Exploring the potential of fine-tuning Large Language Models (LLMs) like Llama2 and StableLM for medical entity extraction. This project focuses on adapting these models using PEFT, Adapter V2, and LoRA techniques to efficiently and accurately extract drug names and adverse side-effects from pharmaceutical texts · GitHub","url":"https://github.com/architkaila/Fine-Tuning-LLMs-for-Medical-Entity-Extraction"},{"title":"GitHub - guilopgar/Medication-Detection-LLM: Detecting Medication Mentions in Social Media Data Using Large Language Model · GitHub","url":"https://github.com/guilopgar/Medication-Detection-LLM"},{"title":"pharma · GitHub Topics","url":"https://github.com/topics/pharma?o=desc&s=updated"},{"title":"LLM Web Scraping: How AI Models Replace Scrapers","url":"https://scrapegraphai.com/blog/llm-web-scraping"},{"title":"llm-pipeline · GitHub Topics · GitHub","url":"https://github.com/topics/llm-pipeline"},{"title":"How to Design Scraping Systems for LLM Training Pipelines | Grepsr","url":"https://www.grepsr.com/blog/llm-data-pipelines-web-scraping-grepsr/"}]

Based on the search results, I found several relevant projects related to your query:

Key Findings

There is an AI-powered pipeline for scraping and tracking pharmaceutical clinical pipeline data with LLM-based change detection, which appears to be available on GitHub under the pharmaceutical-data topic.

The search results also reveal several related projects in this space:

  • Medication Detection: Fine-tuned large language models (LLMs) for the automatic extraction of medication mentions from social media texts, specifically Twitter

  • Medical Entity Extraction: Projects focusing on fine-tuning Large Language Models (LLMs) like Llama2 and StableLM for medical entity extraction

  • LLM Web Scraping for Pharma: Pharmaceutical companies need to track regulatory updates, clinical trial results, and safety information across government sites and medical journals

The first result from the pharmaceutical-data GitHub topic appears to be the most directly relevant to your query, as it explicitly mentions a pipeline for tracking pharmaceutical clinical pipeline data using LLM-based change detection with web scraping capabilities (Python, BeautifulSoup, data pipeline, and LangChain).

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #19] github FDA guidance documents monitor scraper diff regulatory updates tracker python

Web search results for query: "github FDA guidance documents monitor scraper diff regulatory updates tracker python"

Links: [{"title":"GitHub - logiover/fda-data-scraper: FDA data scraper — openFDA drug/device/food recalls, adverse events & 510(k) clearances as JSON/CSV. No API key, no login.","url":"https://github.com/logiover/fda-data-scraper"},{"title":"GitHub - tsbischof/fda: Scraper and aggregator for analysis of FDA 510(k) and related documents. Includes tools for analyzing predicate devices. · GitHub","url":"https://github.com/tsbischof/fda"},{"title":"GitHub - DarpitPatel/OpenFDA: Python code to Scrape data from OpenFDA · GitHub","url":"https://github.com/DarpitPatel/OpenFDA"},{"title":"drugs@fda scraper · GitHub","url":"https://gist.github.com/5997278"},{"title":"GitHub - Tanguy9862/AI-Powered-FDA-Drug-Scraper: Developed a Python-based web scraper leveraging generative AI with LangChain and GPT-4o-mini to extract and classify FDA drug approval data. Processed over 1,770 records, dynamically categorizing medications and treatment areas using LLMs to simplify complex medical information into actionable insights. · GitHub","url":"https://github.com/Tanguy9862/AI-Powered-FDA-Drug-Scraper"},{"title":"GitHub - sheetalkalburgi/web-scraping: Web scraping algorithm for FDA and Health Canada website · GitHub","url":"https://github.com/sheetalkalburgi/web-scraping"},{"title":"GitHub - Norbaeocystin/FDA: scraping and analysis of FDA Drugs Approvals data · GitHub","url":"https://github.com/Norbaeocystin/FDA"},{"title":"GitHub - vshah1016/pharma_scraper: A scraper I made in python to take stocks that are going to have a drug released and have all of its data aggregated to a CSV file where I can sort and make big bux","url":"https://github.com/vshah1016/pharma_scraper"},{"title":"GitHub - jbremz/FDA-Analysis: Scraping and analysis of the (now retired) Drugs@FDA site - with scrapy and pandas","url":"https://github.com/jbremz/FDA-Analysis"}]

I found several relevant GitHub repositories related to FDA data scraping and regulatory monitoring tools. Here's a summary of the results:

FDA Data Scraping Tools

1. FDA Data Scraper (logiover/fda-data-scraper) This tool scrapes U.S. FDA drug, device and food data from the official openFDA API, pulling drug recalls, adverse events, drug labels, device recalls, 510(k) clearances and food recalls into clean JSON format. It supports no-code automation through Zapier, Make, n8n or Pipedream for building regulatory-monitoring workflows.

2. FDA 510(k) Scraper (tsbischof/fda) This project is a scraper and aggregator for analysis of FDA 510(k) and related documents, with tools for analyzing predicate devices.

3. AI-Powered FDA Drug Scraper (Tanguy9862/AI-Powered-FDA-Drug-Scraper) This Python-based web scraper leverages generative AI with LangChain and GPT-4o-mini to extract and classify FDA drug approval data.

4. FDA and Health Canada Web Scraping (sheetalkalburgi/web-scraping) This project uses BeautifulSoup, a Python HTML parser, to extract data from FDA and Health Canada websites.

5. Pharma Scraper (vshah1016/pharma_scraper) This tool scrapes Biopharmcatalyst's PDUFA calendar for FDA approvals and exports the data to a CSV file.

These tools represent various approaches to monitoring and tracking FDA regulatory updates using Python and web scraping techniques.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #20] github Gemini CLI headless cron automation repository example scheduled

Web search results for query: "github Gemini CLI headless cron automation repository example scheduled"

Links: [{"title":"GitHub - addyosmani/gemini-cli-tips: Gemini CLI Tips and Tricks · GitHub","url":"https://github.com/addyosmani/gemini-cli-tips"},{"title":"Automate Your Daily Grind: How Gemini CLI Job Transforms Repetitive Tasks into Smart Workflows · GitHub","url":"https://gist.github.com/HainanZhao/92b43e68850189bfee8f39a2c2581ca6"},{"title":"Headless execution · google-gemini/gemini-cli · Discussion #3215","url":"https://github.com/google-gemini/gemini-cli/discussions/3215"},{"title":"Unleashing Gemini CLI Power in GitHub Actions and Beyond","url":"https://www.leeboonstra.dev/genai/gemini_cli_github_actions/"},{"title":"Automation and triage processes | Gemini CLI","url":"https://geminicli.com/docs/issue-and-pr-automation/"},{"title":"Scheduling Jobs With Gemini CLI and Cron | Testing in Production","url":"https://www.testinginproduction.co/blog/automating-ai-jobs-with-gemini-cli"},{"title":"Gemini CLI Use Cases: 10 Real Workflows for Developers - Gemini CLI All in One","url":"https://geminicli.one/blog/gemini-cli-use-cases-workflows"}]

Based on the search results, here are the key findings for Gemini CLI headless cron automation:

Headless Mode & Scheduled Automation

Gemini CLI in headless mode allows you to provide a prompt via command-line arguments or environment variables with output and exit, making it great for integrating with other tools or triggering AI tasks on a schedule.

Cron Integration Examples

Gemini CLI has a headless mode where the -p flag takes a prompt and returns a response on stdout without an interactive session needed. You can run it from a script with cron to execute prompts as scheduled jobs running in the background while writing reports to your filesystem.

Repository Examples

There are several practical resources available:

  1. Gemini CLI Tips Repository: You can schedule a cronjob that runs a Gemini CLI prompt nightly for tasks like generating reports or cleaning up with AI logic.

  2. Testing in Production: The code for practical cron + Gemini CLI examples is available at github.com/testing-in-production/gemini-jobs.

  3. GitHub Actions Integration: You can combine Gemini CLI's core capabilities with GitHub Actions to create powerful automation, leveraging its headless scripting capabilities since Gemini CLI doesn't have built-in GitHub Actions integration.

Use Cases

Practical uses include calling the playwright MCP, extracting relevant data, classifying it, and integrating it into another system.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[FETCH #1] https://github.com/Q00/data.go.kr-crawling

ASK: Extract: repo description, star count, fork count, primary language, last commit date or last update, file/directory structure, README key contents (what sites/APIs it crawls, libraries used like gevent, requests, BeautifulSoup). Be precise and quote numbers.

Repository Summary: data.go.kr-crawling

Description: "건강정보, 의약품 크롤링, 멀티쓰레딩 gevent" (Health information and pharmaceutical product crawling with multi-threading gevent)

Stars: 4

Forks: 0

Primary Language: Python

Last Commit: 55 commits on development branch (specific date not provided in content)

File Structure

  • apis/ directory
  • async_data_crawler.py
  • go_data_crwaler.py
  • column.py
  • url.py
  • config.py.example
  • requirements.txt
  • README.md
  • .gitignore

Key Details

The project implements a web crawler targeting South Korean government health data portals (data.go.kr). It utilizes gevent for asynchronous multi-threading operations to handle concurrent requests efficiently. The codebase follows a git branching strategy with development and feature branches, merging to master via pull requests with squash commits.

Topics tagged: crawling, gevent, Python, python-lock, threading

[FETCH #2] https://github.com/FDA/openfda

ASK: Extract: repo description, star count, fork count, primary language, last commit date, top-level directory structure, README key contents (Luigi pipelines, which datasets, how to run). Quote numbers precisely.

openFDA Repository Summary

Repository Description: "openFDA is a research project to provide open APIs, raw data downloads, documentation and examples, and a developer community for an important collection of FDA public datasets."

Metrics:

  • Stars: 705
  • Forks: 166
  • Primary Language: Python

Top-Level Directory Structure:

  • api/faers
  • config
  • dependencies
  • openfda
  • schemas
  • scripts
  • Configuration files (Dockerfile, docker-compose.yml, requirements.txt, setup.py)

Key README Contents:

Pipeline Technology: The project uses "Python pipelines written with Luigi for processing public FDA data sets (drugs, foods, medical devices, and other) into a JSON format that can be loaded into Elasticsearch."

Datasets Covered: Drug, food, medical device, and other FDA public datasets, with specific mention of NSDE, CAERS, Substance Data, Device Clearance, Device PMA, and Device Event pipelines.

How to Run: Users can execute docker-compose up to start Elasticsearch and an API container exposing port 8000, with Python containers running data processing pipelines. The documentation notes that "the API container starts right away, it will not serve any data until some or all of the pipelines above have finished running."

Prerequisites: Elasticsearch 7, Python 3.10, Node 16+

[FETCH #3] https://github.com/jbremz/FDA-Analysis

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (scrapy spider structure, pandas analysis, how the Drugs@FDA scrape worked). Quote numbers precisely.

FDA-Analysis Repository Summary

Repository Description: "Scraping and analysis of the (now retired) Drugs@FDA site - with scrapy and pandas"

Repository Statistics:

  • Stars: 5
  • Forks: 2
  • Primary Language: Python

Last Commit Date: Not explicitly stated in the provided content

File Structure: The repository contains:

  • FDA Spider/ A Scrapy project directory for web scraping drug application data
  • masterDrugList2.csv Raw output data from the scraping operation
  • FDA_Data_Analysis.py Pandas-based analysis script
  • Drugs@FDA Analysis.ipynb Jupyter notebook summarizing findings
  • .ipynb_checkpoints/ Notebook checkpoint directory

README Key Contents:

The project involved scraping "over 22,000 different products" from the FDA's Drugs@FDA database at the request of the British Medical Journal to assess data quality and identify missing reports.

Scraping Approach: The FDA Spider directory housed "the main spider was FDASpider," utilizing the Scrapy framework to crawl product application pages.

Analysis Workflow: Analysis of the extracted CSV followed using pandas, with "Drugs@FDA Analysis.ipynb" recommended as the project entry point for understanding methodology and conclusions.

[FETCH #4] https://github.com/logiover/fda-data-scraper

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (which openFDA endpoints, output formats, CLI usage, scheduling/automation notes). Quote numbers precisely.

FDA Data Scraper Repository Summary

Repository Metadata

  • Description: "FDA data scraper — openFDA drug/device/food recalls, adverse events & 510(k) clearances as JSON/CSV"
  • Stars: 0
  • Forks: 0
  • Primary Language: Not specified in content
  • Last Commit Date: Not explicitly stated; repository shows "1 Commit" on master branch

File Structure

The repository contains:

  • /examples directory (with CLI, API, JavaScript, and Python usage examples)
  • .gitignore
  • LICENSE (MIT)
  • README.md

Supported openFDA Endpoints

Nine datasets available:

  1. Drug recalls (enforcement)
  2. Drug adverse events (20M+ records)
  3. Drug labels (258K records)
  4. Device recalls
  5. Device adverse events (24M+ records)
  6. Device 510(k) clearances
  7. Food recalls
  8. Food adverse events
  9. Animal & veterinary adverse events

Output Formats

Supports export as "JSON, CSV, Excel (XLSX), JSONL, XML or HTML"

Key CLI Usage

The documentation provides four implementation methods: Apify Console, Apify CLI, API/curl, and apify-client (JavaScript/Python). Example command structure shows inputs for dataset selection and optional search filtering.

Automation & Scheduling

Features include daily scheduled runs, webhooks on completion, Google Sheets/Excel integration, cloud storage export (S3, GCS), and compatibility with "Zapier, Make, n8n or Pipedream"

[FETCH #5] https://github.com/thp/urlwatch

ASK: Extract: repo description, star count, fork count, primary language, last commit/release date, key features (jobs.yaml, filters, reporters, diff, hooks.py), install/usage. Quote numbers precisely.

urlwatch Repository Summary

Repository Description: "Watch (parts of) webpages and get notified when something changes via e-mail, on your phone or via other means. Highly configurable."

Key Metrics:

  • Stars: 3.1k
  • Forks: 354
  • Primary Language: Python
  • Last Activity: 974 commits on master branch (specific date not provided in page content)

Core Features: The tool monitors webpages for changes and delivers notifications through multiple channels including email and third-party services. It provides unified diffs showing what has changed and includes the modified URL in alerts.

Documentation & Resources:

Topics: automation, monitor, python, webpage

Note: The page content provided does not include specific details about jobs.yaml, filters, reporters, hooks.py implementations, or detailed installation instructions. The repository structure shows these components exist in the codebase (lib/urlwatch directory), but detailed documentation of these features was not captured in the provided web page excerpt.

[FETCH #6] https://github.com/anton-semerenko/pharma-radar

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file/directory structure, README key contents (which sources it monitors: FDA, EMA, etc.; how the LLM agent works; verification policy; Telegram delivery; scheduling; config files). Quote numbers precisely.

Pharma Radar Repository Summary

Repository: anton-semerenko/pharma-radar

Metrics:

  • Stars: 0
  • Forks: 0
  • Primary Language: Python
  • Last Commit: 2 commits on main branch

Description

An autonomous LLM agent that generates daily pharmaceutical intelligence briefings and distributes them via Telegram. The system monitors regulatory approvals, clinical outcomes, market access developments (including Ukraine), and competitive activities across the pharma sector.


Monitored Sources

The agent tracks signals from:

  • Regulators: FDA, EMA, WHO, and Ukrainian authorities (МОЗ, ДЕЦ, Держлікслужба)
  • Trade press: Endpoints News, FiercePharma, STAT News
  • Company investor relations pages and registries

How It Works

Agent Logic: The system uses Claude (Opus-class) in an agentic loop performing: "retrieves broadly, verifies every item against ≥2 independent sources or one official regulatory primary, attributes every number, and states what each development means."

Verification Policy: Items require either two independent sources or one official regulatory primary before inclusion. Single secondary sources are treated as insufficient.

Output Structure: Briefings span four rubrics—Regulatory & Approvals, Clinical & Pipeline, Market/Access & Ukraine, Competitive & Corporate—plus Forward Agenda and Signals to Watch sections. Each item includes 47 sentences with full source attribution.


Delivery & Scheduling

Publishes to a rolling Telegraph page and sends Telegram push notifications. Runs daily via scheduled task (06:00) without human intervention. Uses only Python standard library for the delivery layer.


Repository Structure

├── prompts/system_prompt.md    (agent methodology)
├── config/sources.yaml         (source hierarchy & rubrics)
├── src/deliver.py              (Telegraph + Telegram publishing)
└── examples/sample_digest.md   (sample output)

License: MIT

[FETCH #7] https://github.com/dgtlmoon/changedetection.io

ASK: Extract: repo description, star count, fork count, primary language, last commit/release date, key features (CSS/XPath/jq filters, diff, notification via apprise, playwright/browser steps, schedule/recheck interval, API), docker/pip install instructions. Quote numbers precisely.

changedetection.io Repository Summary

Repository Metrics

  • Stars: 33.5k
  • Forks: 2.0k
  • Primary Language: Python
  • Last Activity: Master branch with 2,448 commits (specific date not provided)

Description

"Best and simplest tool for website change detection, web page monitoring, and website change alerts. Perfect for tracking content changes, price drops, restock alerts, and website defacement monitoring—all for free or enjoy our SaaS plan!"

Key Features

Content Filtering & Detection:

  • CSS Selectors and XPath (1.0 & 2.0) support
  • JSONPath and jq filtering for API monitoring
  • Embedded JSON extraction from HTML pages
  • Visual Selector tool for targeting specific page elements

Monitoring Capabilities:

  • Word, line, and character-level diff visualization
  • Interactive browser steps (login, form filling, button clicks)
  • PDF change tracking
  • Customizable check intervals

Notifications: Apprise library integration supporting Discord, Email, Slack, Telegram, webhooks, and 90+ additional services with Jinja2 templating

Scheduling: Timezone-aware scheduling with day-of-week and time restrictions

Advanced Features:

  • AI-powered change detection and summaries (LiteLLM integration)
  • REST API for programmatic management
  • Chrome extension for easy URL addition
  • Proxy configuration support (including Bright Data integration)

Installation Methods

Docker:

docker compose up -d

Docker Standalone:

docker run -d --restart always -p "127.0.0.1:5000:5000" -v datastore-volume:/datastore --name changedetection.io dgtlmoon/changedetection.io

Pip:

pip3 install changedetection.io
changedetection.io -d /path/to/empty/data/dir -p 5000

Access via browser at http://127.0.0.1:5000

[FETCH #8] https://github.com/huginn/huginn

ASK: Extract: repo description, star count, fork count, primary language, last commit date, key agent types (WebsiteAgent, ChangeDetector, EmailAgent, etc.), how scheduling works. Quote numbers precisely.

Huginn Repository Overview

Repository Description: "Create agents that monitor and act on your behalf. Your agents are standing by!"

Star Count: 49.9k stars

Fork Count: 4.3k forks

Primary Language: Ruby (based on Gemfile, Rails framework, and .rb files throughout)

Last Commit Date: Not explicitly stated in the provided content, though the repository shows "4,134 Commits" on the master branch.

Key Agent Types

The content references these specific agents:

  • WeatherAgent
  • WebsiteAgent (implied through "Scrape websites")
  • EmailAgent (implied through email functionality)
  • TwitterAgent
  • HipChat, FTP, IMAP, Jabber, JIRA, MQTT connectors
  • JavaScript execution agents
  • Location tracking capabilities

Scheduling

The documentation mentions "send digest email with things that you care about at specific times during the day" and references to "Periodically run" commands, but specific scheduling mechanics aren't detailed in this excerpt. The system appears to use event-based propagation: "Huginn's Agents create and consume events, propagating them along a directed graph."

License: MIT

[FETCH #9] https://github.com/jshchnz/claude-code-scheduler

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (how it schedules claude -p, config format, OS support, flags used like --output-format, --allowedTools, --dangerously-skip-permissions). Quote numbers precisely.

claude-code-scheduler Repository Summary

Repo Description: "Put Claude on autopilot. Schedule code reviews, security audits, and anything else - Claude Code runs them automatically, even while you sleep."

Star Count: 510 stars

Fork Count: 37 forks

Primary Language: TypeScript (based on tsconfig.json and src/ directory structure)

Last Commit Date: Not explicitly provided in the content

File Structure

  • .claude-plugin/ - Plugin configuration
  • commands/ - CLI command implementations
  • src/ - Source code
  • skills/scheduler/ - Scheduler skill module
  • examples/ - Usage examples
  • dist/ - Compiled distribution
  • Configuration files: package.json, tsconfig.json, vitest.config.ts

Key README Contents

How Scheduling Works: The plugin registers tasks with native OS schedulers, then executes via claude -p "your command" at scheduled times, logging output to ~/.claude/logs/<task-id>.log.

Supported Flags:

  • --dangerously-skip-permissions - Enables autonomous execution for file edits and commands
  • --output-format - Not explicitly mentioned in README

OS Support: "macOS (launchd), Linux (crontab), Windows (Task Scheduler)"

Config Format: Tasks stored as JSON in .claude/schedules.json (project) or ~/.claude/schedules.json (global), with properties including id, name, trigger (cron expressions), execution (command, timeout, skipPermissions), and optional worktree settings.

Requirements: "Claude Code v1.0.33+"

[FETCH #10] https://github.com/winsw/winsw

ASK: Extract: repo description, star count, fork count, primary language, latest release version and date, last commit date, key XML config options (executable, arguments, onfailure, resetfailure, log mode, startmode, delayedAutoStart), install commands. Quote numbers precisely.

WinSW Repository Summary

Repository Description: "A wrapper executable that can run any executable as a Windows service, in a permissive license."

Engagement Metrics:

  • Stars: 14.3k
  • Forks: 1.7k
  • Primary Language: C#

Release Information: The page indicates stable 2.x releases and 3.x pre-releases are available on GitHub Releases, with NuGet and Maven packages currently supporting 2.x versions. Specific latest version and date details were not displayed in the accessible content.

Latest Activity: The repository shows 841 commits on the v3 branch, though an exact last commit date was not visible in the provided page content.

Key XML Configuration Options: According to the sample configuration file shown, supported options include:

  • <executable> - specifies the application to run
  • <arguments> - command-line parameters
  • <log mode> - logging configuration (example: "roll")
  • <env name> - environment variables

Installation Command: The basic install process uses: winsw install myapp.xml [options]

Additional configuration details are documented in the XML specification file, with samples available in the repository's samples directory.

[FETCH #11] https://github.com/DatGuy1/Windows-Toasts

ASK: Extract: repo description, star count, fork count, primary language, last commit date, install command, usage example, dependencies, Python version support. Quote numbers precisely.

Windows-Toasts Repository Summary

Repository Description: "Python library used to send toast notifications on Windows machines"

Stars: 142

Forks: 9

Primary Language: Python

Last Commit Date: Not explicitly stated in the provided content

Installation Command:

$ python -m pip install windows-toasts

Usage Example:

from windows_toasts import Toast, WindowsToaster
toaster = WindowsToaster('Python')
newToast = Toast()
newToast.text_fields = ['Hello, world!']
newToast.on_activated = lambda _: print('Toast clicked!')
toaster.show_toast(newToast)

Dependencies: The library uses WinRT bindings rather than pywin32. Specific dependency packages are not listed in the extracted content.

Python Version Support: The page indicates "Supported Python versions" with a badge referencing PyPI, but specific version numbers are not provided in the visible content.

Additional Notes: The project supports Windows 10 and 11, with noted limitations compared to pywin32 (duration selection is restricted to short/long rather than specific seconds).

[FETCH #12] https://github.com/GitHub30/win11toast

ASK: Extract: repo description, star count, fork count, primary language, last commit date, install command, key usage examples (toast(), buttons, on_click, image, duration, app_id), dependencies (winsdk/winrt). Quote numbers precisely.

win11toast Repository Summary

Description: "Toast notifications for Windows 10 and 11 based on WinRT"

Stars: 333
Forks: 24
Primary Language: Python
Last Commit Date: Not explicitly displayed in provided content

Installation:

pip install win11toast

Key Usage Examples:

Basic toast:

from win11toast import toast
toast('Hello Python🐍')

With click handler and URL:

toast('Hello Python', 'Click to open url', on_click='https://www.python.org')

With buttons:

toast('Hello', 'Click a button', buttons=['Approve', 'Dismiss', 'Other'])

With image:

toast('Hello', 'Hello from Python', 
      image='https://example.com/image.png')

With duration:

toast('Hello Python🐍', duration='long')

Dependencies:

  • WinRT/Windows SDK (based on WinRT architecture)
  • Acknowledged predecessors: winsdk_toast, Windows-Toasts, MarcAlx/notification.py

License: MIT

[FETCH #13] https://github.com/lorien/awesome-web-scraping

ASK: Extract: repo description, star count, fork count, last commit date, structure (python.md, javascript.md, etc.), sections relevant to change detection, scheduling, and Excel/reporting. Quote numbers precisely.

Repository Analysis: awesome-web-scraping

Repository Metadata

  • Description: "List of libraries, tools and APIs for web scraping and data processing."
  • Stars: 8.1k
  • Forks: 934
  • Last Commit: 640 commits on master branch (specific date not provided in content)

Repository Structure

The repo organizes web scraping resources by programming language and category:

  • python.md - Python packages
  • javascript.md - JavaScript packages
  • php.md - PHP packages
  • ruby.md - Ruby packages
  • golang.md - Go packages
  • cli.md - Command line tools
  • manuals.md - Educational articles and books

Sections Relevant to Change Detection & Scheduling

The content provided does not explicitly detail change detection, scheduling, or monitoring capabilities. The repository appears to be a curated list rather than an active monitoring tool.

Sections Relevant to Excel/Reporting

No Excel or reporting functionality is mentioned in the provided content. This is a documentation repository, not a data processing application.

Additional Resources

  • Links to captcha-solving services
  • Proxy marketplace references
  • Telegram discussion groups for community engagement
  • Contributing guidelines available in CONTRIBUTING.md

[FETCH #14] https://github.com/larsekje/PythonWindowsServices

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (nssm commands, how script is run as service, logging). Quote numbers precisely.

Repository Summary: PythonWindowsServices

Repo Description: "A proof-of-concept on how to run a python script as a windows service using nssm"

Stats:

  • Stars: 1
  • Forks: 0
  • Primary Language: Python
  • Last Commit Date: Not explicitly stated in provided content

File Structure:

  • /logs folder
  • /scripts folder
  • /windows_service folder
  • .gitignore
  • readme.md
  • requirements.txt

Key README Contents:

The repository demonstrates Windows service installation via NSSM (Non-Sucking Service Manager) using three primary commands:

  1. Service installation: nssm install "SERVICE_NAME" "PATH_TO_PYTHON.exe" "PATH_TO_PYTHON_SCRIPT.py"
  2. Optional logging configuration with stdout/stderr file paths
  3. Service startup: nssm start SERVICE_NAME

The proof-of-concept runs a Python script as a Windows service, with logs accessible in the /scripts folder. The setup requires NSSM in system PATH, a configured Python environment with dependencies from requirements.txt, and optionally pre-created log files since NSSM doesn't auto-generate them.

[FETCH #15] https://github.com/HaroldMills/Python-Windows-Service-Example

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (pywin32 service class structure, install/start commands, pitfalls). Quote numbers precisely.

Python-Windows-Service-Example Repository Analysis

Repository Description: "Example of a Windows service implemented in Python"

Metrics:

  • Stars: 21
  • Forks: 11
  • Primary Language: Python
  • Last Commit Date: Not specified in provided content

File Structure: The repository contains:

  • .gitignore
  • LICENSE (MIT)
  • README.md
  • example_service.py
  • example_service.spec (referenced for PyInstaller)

README Key Contents:

Dependencies: "The service should be built in a Python environment that includes the pywin32 and pyinstaller packages."

Build Command: "The command for building the service is: pyinstaller example_service.spec" issued from the repository root.

Installation: Users must run example_service.exe install and example_service.exe start from build\example_service directory.

Important Requirement: "The commands must be issued from a command prompt that was run as administrator."

Removal: To uninstall, execute example_service.exe stop and example_service.exe remove with administrator privileges.

Notable Limitation: The documentation notes that "as of this writing (2016-03-37), PyInstaller supports Python versions only through 3.5"—indicating this project may require updates for modern Python compatibility.

[FETCH #16] https://github.com/Bwhiz/Auto-Excel-Reports

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (openpyxl report generation, GitHub Actions cron, SMTP email). Quote numbers precisely.

Auto-Excel-Reports Repository Summary

Repository Description: "Scripts and workflows to automate the generation and distribution of Excel reports using Python's openpyxl library" with GitHub Actions enabling "scheduled and event-triggered report generation and email distribution."

Metrics:

  • Stars: 1
  • Forks: 0
  • Primary Language: Python

File Structure: The repository contains .github/workflows/ (automation configurations), assets/ (sample images), report_script.py, auto_mail.py, requirements.txt, .gitignore, and README.md.

Key README Features:

  1. Report Generation: Utilizes openpyxl for creating and formatting Excel workbooks with customizable content.

  2. GitHub Actions Automation: Implements cron-based scheduling with example workflow showing schedule: - cron: '0 0 * * *' for daily execution.

  3. Email Distribution: Integrates SMTP functionality via auto_mail.py using environment variables for sender credentials and recipient addresses.

  4. Security: Leverages GitHub Secrets to "securely handle sensitive information like email credentials."

Note: The page displays current activity but does not specify the exact last commit date in accessible text.

[SEARCH #21] "nedrug.mfds.go.kr" github python selenium requests 크롤링 프로젝트

Web search results for query: ""nedrug.mfds.go.kr" github python selenium requests 크롤링 프로젝트"

Links: [{"title":"[Python] Selenium으로 웹사이트 크롤링하기(2) — 웹 사이트 제어하기 | by Sarah Na | Medium","url":"https://2island.medium.com/python-selenium%EC%9C%BC%EB%A1%9C-%EC%9B%B9%EC%82%AC%EC%9D%B4%ED%8A%B8-%ED%81%AC%EB%A1%A4%EB%A7%81%ED%95%98%EA%B8%B0-2-%EC%9B%B9-%EC%82%AC%EC%9D%B4%ED%8A%B8-%EC%A0%9C%EC%96%B4%ED%95%B4%EB%B3%B4%EA%B8%B0-1ffc5e05179d"},{"title":"Mediteam.us 개발 - Python & Selenium을 이용한 구글검색 크롤링 — Steemit","url":"https://steemit.com/kr/@junn/mediteam-us-python-and-selenium"},{"title":"Python Selenium 사용법 [파이썬 셀레늄 사용법, 크롤링]","url":"https://greeksharifa.github.io/references/2020/10/30/python-selenium-usage/"},{"title":"GPT가 알려주는데로 크롤링 만들기 (with python, selenium) - Summers Blog","url":"https://sunmerrr.github.io/other/crawling-1/"},{"title":"Selenium으로 네이버 연극 데이터 크롤링하기 with Python","url":"https://teamlab.github.io/jekyllDecent/blog/crawling%20with%20python/Selenium%EC%9C%BC%EB%A1%9C-%EB%84%A4%EC%9D%B4%EB%B2%84-%EC%97%B0%EA%B7%B9-%EB%8D%B0%EC%9D%B4%ED%84%B0-%ED%81%AC%EB%A1%A4%EB%A7%81%ED%95%98%EA%B8%B0-with-Python"},{"title":"[python] Selenium을 이용한 웹 크롤링 - 간단 사용법 및 예제 | JaeSeoKim's Blog","url":"https://jaeseokim.dev/Python/python-Selenium%EC%9D%84-%EC%9D%B4%EC%9A%A9%ED%95%9C-%EC%9B%B9-%ED%81%AC%EB%A1%A4%EB%A7%81-%EA%B0%84%EB%8B%A8-%EC%82%AC%EC%9A%A9%EB%B2%95-%EB%B0%8F-%EC%98%88%EC%A0%9C/"}]

Based on the search results, I found information about web crawling projects using Python with Selenium and Requests, though the specific "nedrug.mfds.go.kr" GitHub project wasn't directly returned.

However, the search results do provide useful context about Python crawling projects with these technologies:

Selenium can authenticate and pass cookies to Python requests for faster crawling, which is a common pattern used in web crawling projects.

Selenium is a Python package that is an essential library for web crawling alongside BeautifulSoup4.

Selenium is developed for web automation testing and uses an actual browser, allowing observation of DOM changes, while the requests library simply sends requests and retrieves the response values from the server for crawling.

If you're looking for a specific GitHub project related to "nedrug.mfds.go.kr" (the Korean Ministry of Food and Drug Safety database), the search results didn't return that specific repository. You might want to search directly on GitHub with those terms, or provide more specific details about what you're looking for regarding that crawling project.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #22] github Korea MFDS drug approval scraper "nedrug" OR "mfds" python

Web search results for query: "github Korea MFDS drug approval scraper "nedrug" OR "mfds" python"

Links: [{"title":"https:/nedrug.mfds.go.kr - 식품의약품안전처","url":"https://nedrug.mfds.go.kr/"},{"title":"OPENAPI Detail | PUBLIC DATA PORTAL","url":"https://www.data.go.kr/en/data/15117134/openapi.do"},{"title":"How to launch a clinical trial in South Korea: Investigational New Drug Application Process","url":"https://www.precisionformedicine.com/blog/how-to-launch-a-clinical-trial-in-south-korea-investigational-new-drug-application-process"},{"title":"Ministry of Food and Drug Safety>Our Works>Drugs>GIFT | Ministry of Food and Drug Safety","url":"https://www.mfds.go.kr/eng/wpge/m_1176/de011009l001.do"},{"title":"MFDS Drug Safety Korea - 식품의약품안전처","url":"https://nedrug.mfds.go.kr/eng/index"},{"title":"Welcome to the Ministry of Food and Drug Safety | Ministry of Food and Drug Safety","url":"https://www.mfds.go.kr/eng/index.do"},{"title":"Ministry of Food and Drug Safety>Our Works>Drugs>Approval Process | Ministry of Food and Drug Safety","url":"https://www.mfds.go.kr/eng/wpge/m_17/denofile.do"},{"title":"GitHub - Tanguy9862/AI-Powered-FDA-Drug-Scraper: Developed a Python-based web scraper leveraging generative AI with LangChain and GPT-4o-mini to extract and classify FDA drug approval data. Processed over 1,770 records, dynamically categorizing medications and treatment areas using LLMs to simplify complex medical information into actionable insights. · GitHub","url":"https://github.com/Tanguy9862/AI-Powered-FDA-Drug-Scraper"},{"title":"GitHub - vshah1016/pharma_scraper: A scraper I made in python to take stocks that are going to have a drug released and have all of its data aggregated to a CSV file where I can sort and make big bux","url":"https://github.com/vshah1016/pharma_scraper"},{"title":"Kim Gang-lip","url":"https://en.wikipedia.org/wiki/Kim_Gang-lip"}]

Based on the search results, here's what I found:

The search results show information about the MFDS (Ministry of Food and Drug Safety) in South Korea and their drug approval portal at nedrug.mfds.go.kr, but they don't contain specific GitHub repositories dedicated to scraping Korean MFDS drug approval data.

What the results do show:

  • The MFDS drug information can be searched in real time through an integrated search at nedrug.mfds.go.kr/searchDrug
  • The MFDS Drug Information Portal at https://nedrug.mfds.go.kr is used for submission of IND applications

The search results include examples of FDA drug approval scrapers on GitHub (like the AI-Powered-FDA-Drug-Scraper and pharma_scraper projects), but there don't appear to be specific results for a Python scraper targeting the Korean MFDS/nedrug system in these results.

If you're looking for a specific GitHub repository for MFDS scraping, you may want to search directly on GitHub using the terms "MFDS scraper" or "nedrug python" to find active projects in that space.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #23] github openFDA python package pypi client library "openfda" wrapper

Web search results for query: "github openFDA python package pypi client library "openfda" wrapper"

Links: [{"title":"GitHub - rOpenHealth/openfda: Convenient access to the OpenFDA API · GitHub","url":"https://github.com/rOpenHealth/openfda"},{"title":"openfda · GitHub Topics · GitHub","url":"https://github.com/topics/openfda?l=r&o=desc&s=updated"},{"title":"OpenFDA Drug Intelligence + AI Severity API in Python · Apify","url":"https://apify.com/benthepythondev/openfda-drug-intelligence/api/python"},{"title":"github · PyPI","url":"https://pypi.org/project/github/"},{"title":"Food and Drug Administration · GitHub","url":"https://github.com/fda"},{"title":"pyfda · PyPI","url":"https://pypi.org/project/pyfda/"},{"title":"openfda/openfda/faers/pipeline.py at master · FDA/openfda","url":"https://github.com/FDA/openfda/blob/master/openfda/faers/pipeline.py"},{"title":"GitHub - realpython/list-of-python-api-wrappers: List of Python API Wrappers and Libraries · GitHub","url":"https://github.com/realpython/list-of-python-api-wrappers"},{"title":"github.com","url":"https://github.com/FDA/openfda"},{"title":"github.com","url":"https://github.com/roivant/openfda"}]

Based on the search results, here's what I found regarding openFDA Python packages:

Official openFDA Project

The official openFDA is an FDA project to provide open APIs, raw data downloads, documentation and examples, and a developer community for an important collection of FDA public datasets. The repository contains Python pipelines written with Luigi for processing public FDA data sets (drugs, foods, medical devices, and other) into a JSON format that can be loaded into Elasticsearch.

The official GitHub repository is available at: https://github.com/FDA/openfda

Third-Party Python Wrappers

There are also community-created Python wrappers for the openFDA API:

  1. rOpenHealth/openfda (R package, not Python): While not a Python package, this is worth noting as it provides convenient access to the OpenFDA API and offers simple helpers for accessing the API with the jsonlite and magrittr packages.

  2. LLM-powered wrappers: There are LLM-powered wrappers for the OpenFDA API that provide access to all publicly available FDA data using a natural language interface.

The search results show that while there are various tools and wrappers for accessing openFDA data, the primary official implementation is the FDA's own repository rather than a published PyPI package specifically named "openfda."

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[FETCH #17] https://github.com/testing-in-production/gemini-jobs

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (how gemini CLI headless -p is run from cron, scripts, output handling). Quote numbers precisely.

The server returned HTTP 404 Not Found.

The response body was not retrieved. If this URL requires authentication, use an authenticated tool (e.g. gh for GitHub, or an MCP-provided fetch tool) instead of WebFetch.

[SEARCH #24] github python scraper diff snapshot "new" "changed" "removed" rows compare previous run excel report sqlite

Web search results for query: "github python scraper diff snapshot "new" "changed" "removed" rows compare previous run excel report sqlite"

Links: [{"title":"GitHub - ecprice/newsdiffs: Automatic scraper that tracks changes in news articles over time. · GitHub","url":"https://github.com/ecprice/newsdiffs"},{"title":"GitHub - larsyencken/csvdiff: Generate a diff between two tabular datasets expressed in CSV files.","url":"https://github.com/larsyencken/csvdiff"},{"title":"python-scraper · GitHub Topics · GitHub","url":"https://github.com/topics/python-scraper"},{"title":"Git scraping: track changes over time by scraping to a Git repository","url":"https://simonwillison.net/2020/Oct/9/git-scraping/"},{"title":"How to Scrape GitHub Data Repository With Python","url":"https://www.scraperapi.com/web-scraping/github/"},{"title":"Releases · firecrawl/firecrawl","url":"1049d38a30"}]

Based on the web search results, I found several relevant resources for your query about Python scrapers with diff/snapshot capabilities and data comparison:

Key Findings:

1. csvdiff - CSV Comparison Tool

The csvdiff tool can generate diffs showing "rows removed," "rows added," and "rows changed" with detailed field-level differences. This tool is particularly useful for comparing tabular data and can output results in JSON format with categories for added, changed, and removed rows.

2. Git Scraping Approach

Git scrapers can grab data periodically, commit it to a repository if changed, creating a commit log of changes to information over time. A tool called git-history was released to help analyze data collected using this technique.

3. News Article Scraper (newsdiffs)

The newsdiffs project demonstrates automatic scraping that tracks changes in articles, where running the scraper again after a time interval and comparing snapshots displays any associated changes.

4. GitHub Data Scraping

Various Python scrapers exist for extracting GitHub repository data and can be extended to track snapshots of data over time for comparison purposes.

For your specific use case involving Excel reports and SQLite databases, you might want to combine csvdiff for row-level comparison with a storage solution (SQLite) and reporting tool (Excel export) to track changes between scraper runs.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #25] github watchdog script restart windows service toast notification powershell BurntToast

Web search results for query: "github watchdog script restart windows service toast notification powershell BurntToast"

Links: [{"title":"Display toast notifications with PowerShell's BurntToast module | PDQ","url":"https://www.pdq.com/blog/display-toast-notifications-with-powershell-burnt-toast-module/"},{"title":"BurntToast Button to reboot the computer · Windos/BurntToast · Discussion #140","url":"https://github.com/Windos/BurntToast/discussions/140"},{"title":"GitHub - NakedPowerShell/BurntToast: Module for creating and displaying Toast Notifications on Microsoft Windows 10. · GitHub","url":"https://github.com/NakedPowerShell/BurntToast"},{"title":"Win10 Toast Notification with Schedule buttton to pick a time · Windos/BurntToast · Discussion #179","url":"https://github.com/Windos/BurntToast/discussions/179"},{"title":"Monitoring with PowerShell: Notifying users of Windows Updates","url":"https://www.cyberdrain.com/monitoring-with-powershell-notifying-users-of-windows-updates/"},{"title":"GitHub - Windos/BurntToast: Module for creating and displaying Toast Notifications on Microsoft Windows 10. · GitHub","url":"https://github.com/Windos/BurntToast"},{"title":"Powershell : burnt toast notification - PowerShell Help - PowerShell Forums","url":"https://forums.powershell.org/t/powershell-burnt-toast-notification/24222"},{"title":"Reboot Notifications with BurntToast: A Simple Guide | Joshua Dearing","url":"https://www.dearing.dev/posts/Reboot-Notifications-with-BurntToast-A-Simple-Guide/"},{"title":"github.com","url":"https://github.com/Badgerati/Hook"},{"title":"github.com","url":"https://github.com/michalzobec/autorunsalerts"}]

I found several relevant results for your search. Here's what I discovered:

Key Resources Found:

Main BurntToast Repository: BurntToast is a PowerShell Module for displaying Toast Notifications on Windows 10 and Windows Server 2019 and above. You can find the official repository at github.com/Windos/BurntToast.

BurntToast Capabilities: BurntToast is a PowerShell module that lets you create custom Windows toast notifications — perfect for script completion alerts, Pomodoro timers, or nudging users to reboot.

Service Monitoring Example: The search results included an example using the Hook PowerShell module (github.com/Badgerati/Hook) that demonstrates monitoring a service and sending a pop-up if the service stops, such as checking the MongoDB service status with BurntToast notifications.

Practical Watchdog Implementation: There's also a project called AutorunsAlerts (github.com/michalzobec/autorunsalerts) that demonstrates a watchdog pattern, where one scheduled task runs as system every 60 minutes to scan for changes, while another task runs in user context every 15 minutes to send toast alerts.

These resources show how to create watchdog-style scripts that monitor Windows services and send toast notifications using BurntToast and PowerShell.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #26] github Codex CLI "codex exec" headless cron automation pipeline example

Web search results for query: "github Codex CLI "codex exec" headless cron automation pipeline example"

Links: [{"title":"OpenAI Codex CLI: Complete Getting Started Guide","url":"https://www.deployhq.com/blog/getting-started-with-openai-codex-cli-ai-powered-code-generation-from-your-terminal"},{"title":"Codex CLI Headless and Batch Mode: Non-Interactive Automation Guide | Codex Knowledge Base","url":"https://codex.danielvaughan.com/2026/04/18/codex-cli-headless-batch-mode-automation/"},{"title":"Codex Automations as Lightweight CI: When Scheduled Agents Replace Your Pipeline | Codex Knowledge Base","url":"https://codex.danielvaughan.com/2026/07/19/codex-automations-lightweight-ci-scheduled-agents-codex-exec-github-actions/"},{"title":"Codex CLI Automations and Scheduled Tasks: Background Agent Workflows | Codex Knowledge Base","url":"https://codex.danielvaughan.com/2026/03/27/codex-cli-automations-scheduled-tasks/"},{"title":"Codex CLI for CI/CD: codex exec, Non-Interactive Mode and Pipeline Integration | Codex Knowledge Base","url":"https://codex.danielvaughan.com/2026/03/26/codex-cli-cicd-non-interactive/"},{"title":"Codex CLI Automation: 3 Workflow Patterns for GitHub Actions, Cron & CI - SmartScope","url":"https://smartscope.blog/en/generative-ai/chatgpt/codex-cli-automation-workflow-patterns/"},{"title":"Codex Exec in CI: The Practical Guide to Headless OpenAI Agents - Developers Digest","url":"https://www.developersdigest.tech/blog/codex-exec-ci-headless-guide"},{"title":"How to Set Up a CI/CD Pipeline with Codex | Codexlog","url":"https://codexlog.dev/guides/tasks/setup-ci-cd-pipeline/"}]

I found relevant information about GitHub Codex CLI's codex exec command for headless automation and pipeline integration. Here are the key findings:

Overview

The codex exec command runs Codex without the interactive terminal UI, takes a prompt, executes it, and exits — making it ideal for CI/CD pipelines, Git hooks, cron jobs, and scripted automation.

How It Works

When you invoke codex exec, Codex starts a single agent session, executes the task to completion, streams progress to stderr, writes the final agent message to stdout, and exits. There is no approval prompt — the agent runs autonomously under whatever sandbox policy you configure.

Cron Job Example

A daily triage workflow example uses codex exec --full-auto in a GitHub Actions scheduled job (cron: '0 8 * * 1-5') to review open GitHub issues older than 14 days.

Batch Processing Pipeline Example

A CSV-driven batch automation pattern uses a bash loop with codex exec --full-auto --ephemeral to process multiple files from a task list, running separate agent sessions for each task.

Key Flags for Automation

The default sandbox for codex exec is read-only, suitable for analysis tasks. Use --sandbox workspace-write in new scripts for file modifications, and --ask-for-approval never for unattended runs.

The search results include practical examples from official Codex documentation and community guides showing how to integrate codex exec into production automation workflows.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[FETCH #18] https://github.com/coderxio/OpenFDA

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (which openFDA datasets, database used, docker compose, how to run). Quote numbers precisely.

OpenFDA Repository Summary

Repo Description: "Python scripts for capturing OpenFDA data in a database."

Star Count: 3 stars

Fork Count: 3 forks

Primary Language: Python

Last Commit Date: Not explicitly stated in the provided content

File Structure

The repository contains:

  • openfda/ folder
  • .gitignore
  • README.md
  • docker-compose.override.yml
  • docker-compose.yml

README Key Contents

OpenFDA Dataset: The project uses the drug NDC dataset from "https://open.fda.gov/apis/drug/ndc/download/" with instructions to unzip and rename it as 'drug-ndc.json' before placing in the ./data/ folder.

Database & Docker: The README provides two execution methods:

  1. Virtual environment approach: Create venv, install requirements, then run python app/load_data.py and python app/serve_data.py

  2. Docker approach: "First run to load DB: docker-compose up --build" followed by production deployment using a separate compose file, with cleanup via docker-compose down -v

The setup uses CherryPy as the server framework.

[FETCH #19] https://github.com/topics/pharmaceutical-data

ASK: List every repository shown with owner/name, description, stars, language, last updated. Especially find the one described as AI-powered pipeline for scraping and tracking pharmaceutical clinical pipeline data with LLM-based change detection.

Pharmaceutical Data Repositories on GitHub

Here are the 7 public repositories tagged with pharmaceutical-data:

  1. kawsarlog/AmerisourceBergen (4 stars)

    • Python | Updated Aug 5, 2023
    • "automates the extraction of product pricing details from the AmerisourceBergen website"
  2. mrueda/nomenclator-delta (1 star)

    • Python | Updated Aug 9, 2026
    • Monitors monthly changes in medicines and health products from Spain's health ministry
  3. MohammedAhmed-01/DataDoseProject (0 stars)

    • Jupyter Notebook | Updated Mar 22, 2026
    • "End-to-end pharmaceutical data pipeline for active ingredient validation, OpenFDA label enrichment, and systematic drugdrug interaction detection"
  4. Dagiayy/kara-medical-telegram-data-platform (0 stars)

    • Python | Updated Aug 23, 2026
    • Manages Ethiopian medical/pharmaceutical data via Telegram scraping, PostgreSQL storage, and Dagster orchestration
  5. bdmorris238/pharmaco-database-project (0 stars)

    • PLpgSQL | Updated Aug 15, 2025
    • Relational and data warehouse design with SQL scripts for pharmaceutical market analysis
  6. Mzands2622/Zanalytix (0 stars)

    • Python | Updated Mar 9, 2026
    • "AI-powered pipeline for scraping and tracking pharmaceutical clinical pipeline data with LLM-based change detection"
  7. khushihajiyani-dotcom/drug-spending-analysis (0 stars)

    • SQLite | Updated May 3, 2026
    • Analyzes drug spending across Canadian provinces (2020-2024)

[FETCH #20] https://github.com/ysfchn/toasted

ASK: Extract: repo description, star count, fork count, primary language, last commit date, install command, key usage example, Python version support, dependencies. Quote numbers precisely.

Toasted Repository Summary

Description: "Rich toast notifications library for Windows (10/11) in Python, supports all elements provided by Windows."

Stars: 31

Forks: 2

Primary Language: Python

Last Commit Date: Not explicitly stated in the provided content

Install Command:

python -m pip install toasted

Key Usage Example:

from toasted import Toast, Progress, Text
import asyncio

async def main():
    toast = Toast()
    toast.elements = [
        Text("File downloader"),
        Progress(value="{value}", status="Downloading files...")
    ]
    await toast.show(dict(value=75/100))

Python Version Support: Not specified in the provided content

Dependencies: The repository includes requirements.lock and requirements-dev.lock files, but specific dependency versions are not detailed in the excerpt provided.

Notable Feature: The library uniquely supports "every notification element provided by Windows, like images, select, inputs" rather than limiting to text-only notifications.

[FETCH #21] https://drew.tech/posts/claude-code-as-a-cron-job

ASK: Extract the exact shell/cron commands and claude flags used (e.g. -p, --output-format, --allowedTools, --dangerously-skip-permissions, --max-turns), how output is handled, and any gotchas mentioned (PATH, auth, env vars).

Claude Code Flags and Commands

Shell Command Used

The article demonstrates this invocation:

claude \
  --dangerously-skip-permissions \
  --output-format json \
  --json-schema "$(cat /tmp/schema.json)" \
  -p "$(cat /tmp/prompt.txt)" \
  > /tmp/output.json

Flags Documented

Flag Purpose
--dangerously-skip-permissions Bypasses permission prompts
--output-format json Returns structured JSON output
--json-schema Validates output against provided schema
-p Accepts prompt from file or stdin

Output Handling

The workflow captures output by:

  • Redirecting stdout to /tmp/output.json
  • Reading the file with sandbox.readFileToBuffer()
  • Parsing JSON and extracting the structured_output field
  • Validating against a Zod schema

Gotchas Mentioned

Authentication state: The article emphasizes that "snapshots preserve auth state," meaning MCPs and integrations must be pre-configured once in the snapshot before automation begins.

No PATH or environment variable issues are explicitly mentioned in the content provided.

Additional Context

The author notes that Claude Code is invoked as a subprocess within a Vercel Sandbox that has been pre-booted from a snapshot containing MCPs and custom skills already installed.

[FETCH #22] https://github.com/suriyadeepan/WebScraping-for-Healthcare

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (which sites scraped, libs). Quote numbers precisely.

Repository Summary

Repository: WebScraping-for-Healthcare

Description: "Scraping the internet for extracting healthcare and pharma data."

Stats:

  • Stars: 8
  • Forks: 1
  • Primary Language: Python

Last Commit: 44 total commits on main branch (specific date not displayed in content)

File Structure

├── data/
├── mhra/
├── notebooks/
├── phscrape/
├── twitter/
├── .gitignore
├── LICENSE (GPL-3.0)
├── README.md
├── requirements.txt
└── tests.py

Key Contents

Data Sources Scraped:

  • Drug Bank (pharmaceutical database)
  • ClinicalTrials.gov (clinical trial information)
  • COVID-19 API
  • Twitter (#remdesivir hashtag)
  • EMC (Electronic Medicines Compendium) - SMPC/PIL extraction
  • HPRA (Health Products Regulatory Authority)
  • MHRA (UK Medicines and Healthcare Products Regulatory Agency)

Primary Libraries Used: The phscrape module provides functions like drugbank.fetch(), clinicaltrials.fetch(), emc.crawl_k(), and hpra.crawl_k() for automated data extraction across these sources.

[FETCH #23] https://github.com/winsw/winsw/blob/v3/docs/xml-config-file.md

ASK: Extract the XML config elements for: onfailure (actions restart/reboot/none, delay), resetfailure, startmode (Automatic, Manual, Boot, System) and delayedAutoStart, log mode options (append, reset, roll, roll-by-size, roll-by-time, none), stopexecutable/stoparguments, env, workingdirectory, priority, serviceaccount. Give exact element names and example snippets.

WinSW XML Configuration Elements

onfailure

Element: <onfailure>

Actions: restart, reboot, none

<onfailure action="restart" delay="10 sec"/>
<onfailure action="restart" delay="20 sec"/>
<onfailure action="reboot" />

The delay attribute supports: sec/secs/min/mins/hour/hours/day/days.

resetfailure

Element: <resetfailure>

<resetfailure>1 hour</resetfailure>

"Controls the timing in which Windows SCM resets the failure count." Defaults to 1 day.

startmode

Element: <startmode>

Supported values: Automatic, Manual

<startmode>Automatic</startmode>

Default is Automatic.

delayedAutoStart

Element: <delayedAutoStart>

<delayedAutoStart>true</delayedAutoStart>

"Enables the delayed start mode if the Automatic start mode is defined."

log

Element: <log>

Mode options: append (default), reset, ignore, roll

<log mode="roll"></log>

stopexecutable/stoparguments

Elements: <stopexecutable>, <stoparguments>

<executable>catalina.sh</executable>
<startarguments>jpda run</startarguments>

<stopexecutable>catalina.sh</stopexecutable>
<stoparguments>stop</stoparguments>

env

Element: <env>

<env name="HOME" value="c:\\abc" />

workingdirectory

Element: <workingdirectory>

<workingdirectory>C:\\application</workingdirectory>

priority

Element: <priority>

Values: idle, belownormal, normal, abovenormal, high, realtime

<priority>idle</priority>

serviceaccount

Element: <serviceaccount>

<serviceaccount>
  <username>DomainName\\UserName</username>
  <password>Pa55w0rd</password>
  <allowservicelogon>true</allowservicelogon>
</serviceaccount>

For Group Managed Service Accounts, append $ to username and omit <password>.

[FETCH #24] https://nssm.cc/usage

ASK: Extract NSSM CLI usage: install, set AppDirectory, AppParameters, AppStdout, AppStderr, AppExit actions (Restart/Ignore/Exit), AppRestartDelay, AppThrottle, AppRotateFiles, Start (SERVICE_AUTO_START, SERVICE_DELAYED_AUTO_START), and example commands.

NSSM CLI Usage Guide

Installation Command

nssm install <servicename> <application> [<options>]

The basic service installation bypasses the GUI, allowing command-line configuration.

Key Configuration Commands

Application Path & Directory:

nssm set <servicename> Application C:\path\to\app.exe
nssm set <servicename> AppDirectory C:\startup\directory
nssm set <servicename> AppParameters argument1 argument2

I/O Redirection:

nssm set <servicename> AppStdout C:\path\to\output.log
nssm set <servicename> AppStderr C:\path\to\error.log

Service Start Behavior:

nssm set <servicename> Start SERVICE_AUTO_START
nssm set <servicename> Start SERVICE_DELAYED_AUTO_START

Exit & Restart Configuration

Exit Actions define responses when the application terminates:

  • Restart — automatically relaunch the service
  • Ignore — leave service stopped
  • Exit — stop the service

Set via registry: HKLM\System\CurrentControlSet\Services\<servicename>\Parameters\AppExit

Restart Throttling prevents CPU loops:

nssm set <servicename> AppThrottle 1500

Delays restarts if the application exits within the threshold milliseconds.

Restart Delay enforces mandatory intervals between restarts:

nssm set <servicename> AppRestartDelay 3000

File Rotation

Enable automatic log rotation:

nssm set <servicename> AppRotateFiles 1

Service Removal

nssm remove <servicename> confirm

Example: Complete Service Configuration

nssm install MyService "C:\Program Files\app.exe"
nssm set MyService AppDirectory C:\Program Files
nssm set MyService AppParameters --config settings.ini
nssm set MyService AppStdout C:\logs\output.log
nssm set MyService AppStderr C:\logs\error.log
nssm set MyService Start SERVICE_AUTO_START
nssm set MyService AppThrottle 2000
nssm set MyService AppRestartDelay 5000

[FETCH #25] https://code.claude.com/docs/en/headless

ASK: Extract: exact flags for headless/print mode (-p / --print, --output-format text|json|stream-json, --json-schema or structured output, --allowedTools, --disallowedTools, --permission-mode, --dangerously-skip-permissions, --max-turns, --model, --append-system-prompt, --continue/--resume, --verbose), stdin piping example, exit codes, and any Windows/PowerShell notes. Quote exact flag names.

Documentation Index

Fetch the complete documentation index at: https://code.claude.com/docs/llms.txt Use this file to discover all available pages before exploring further.

Run Claude Code programmatically

Use the Agent SDK to run Claude Code programmatically from the CLI, Python, or TypeScript.

The Agent SDK gives you the same tools, agent loop, and context management that power Claude Code. It's available as a CLI for scripts and CI/CD, or as Python and TypeScript packages for full programmatic control.

To run Claude Code in non-interactive mode, pass -p with your prompt and the CLI options you need:

claude -p "Find and fix the bug in auth.py" --allowedTools "Read,Edit,Bash"

This page covers using the Agent SDK via the CLI (claude -p). For the Python and TypeScript SDK packages with structured outputs, tool approval callbacks, and native message objects, see the full Agent SDK documentation.

Basic usage

Add the -p (or --print) flag to any claude command to run it non-interactively. Not every CLI option combines with -p. Claude Code rejects --bg, and rejects --cloud with a task description, with an error naming the conflict; --cloud with a session ID and -p instead queues a message into that cloud session and exits. Options you'll combine with -p often include:

This example asks Claude a question about your codebase and prints the response:

claude -p "What does the auth module do?"

Claude Code exits with code 0 on success and a non-zero code when the run fails, so your scripts can branch on the exit status. If you pass an invalid flag, Claude Code reports the error to stderr before the run starts. When a failure happens inside the run, such as missing authentication, Claude Code prints the failure as the result on stdout.

Start faster with bare mode

Add --bare to reduce startup time by skipping auto-discovery of hooks, skills, custom commands, subagents, plugins, MCP servers, auto memory, and CLAUDE.md. Without it, claude -p loads the same context an interactive session would, including anything configured in the working directory or ~/.claude.

Bare mode is useful for CI and scripts where you need the same result on every machine. A hook in a teammate's ~/.claude or an MCP server in the project's .mcp.json won't run, because bare mode never reads them. A directory you name with --add-dir is a partial exception: bare mode loads skills from its .claude/skills/ folder, but still skips its .claude/commands/ and .claude/agents/ folders. Skills from additional directories covers what does and doesn't load.

Without --bare, a -p session runs the hooks in a project's .claude/settings.json and connects the servers in its .mcp.json, even in a folder you've never trusted. A -p session shows no workspace trust dialog and no per-server approval prompt. What runs before you trust a folder covers each kind of repository content under -p and how to keep it out.

This example runs a one-off summarize task in bare mode and pre-approves the Read tool so the call completes without a permission prompt. Set ANTHROPIC_API_KEY before running it, because bare mode doesn't use your subscription login:

claude --bare -p "Summarize README.md" --allowedTools "Read"

In bare mode, Claude Code never reads OAuth credentials or the system keychain. For the Anthropic API, set ANTHROPIC_API_KEY in the environment, with a key created in the Claude Console, or supply an apiKeyHelper in the --settings JSON. Amazon Bedrock, Google Cloud's Agent Platform, and Microsoft Foundry continue to read their own provider credentials as usual.

In bare mode Claude has access to the Bash, file read, and file edit tools. Pass any context you need with a flag:

To load Use
System prompt additions --append-system-prompt, --append-system-prompt-file
Settings --settings <file-or-json>
MCP servers --mcp-config <file-or-json>
Custom agents --agents <json>
A plugin --plugin-dir <path>, --plugin-url <url>
`--bare` is the recommended mode for scripted and SDK calls, and will become the default for `-p` in a future release.

Background tasks at exit

If Claude starts a background Bash task during a claude -p run, for example a dev server or a watch build, that shell is terminated about five seconds after Claude has returned its final result and stdin has closed. The grace period lets a task that finishes right after the result still deliver its output. Before v2.1.163, a never-exiting background process would hold the claude -p invocation open indefinitely.

Background subagents and workflows are exempt from the five-second grace because their result is part of the final output, so claude -p waits for them to complete. From v2.1.182, that wait is capped at ten minutes of continuous idle waiting by default, so a stuck background agent can't hold the process open indefinitely. Adjust the cap with CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS, or set it to 0 to wait without a limit.

Stop a run with SIGTERM

If you stop a claude -p run with SIGTERM, for example with kill or from a process supervisor, Claude Code exits with code 143. Claude Code leaves the turn that was in progress unfinished and records no result for it. To end the turn instead, send SIGINT, or call the Agent SDK's interrupt(), before you stop the process.

On SIGTERM, Claude Code terminates the process tree of any Bash command that is still running. Claude Code then runs SessionEnd hooks and exits. While exiting, Claude Code starts no new tool call, sends no new model request, and runs no hook other than SessionEnd. If the run was in the middle of a command or waiting on a permission prompt when the signal arrived, Claude Code handles that step as follows:

  • Running a command: Claude Code records the command as killed in the session.
  • Waiting for an answer to a permission prompt: if you send SIGTERM to the process, Claude Code leaves the prompt unanswered. If your program closes the session through the Agent SDK, the SDK ends Claude Code's input before sending any signal, and Claude Code cancels the prompt as soon as the input ends.

When you resume the session, Claude Code continues the turn that SIGTERM left unfinished.

Examples

These examples highlight common CLI patterns. Where a command names a file such as auth.py or build-error.txt, substitute a file from your own project. In CI or other scripted environments, add --bare so Claude Code starts without loading the host's hooks, plugins, auto memory, or CLAUDE.md.

Pipe data through Claude

Non-interactive mode reads stdin, so you can pipe data in and redirect the response out like any other command-line tool.

This example pipes a build log into Claude and writes the explanation to a file:

cat build-error.txt | claude -p 'concisely explain the root cause of this build error' > output.txt

With --output-format json, the response payload includes total_cost_usd and a per-model cost breakdown, so scripted callers can track spend per invocation without consulting the usage dashboard. Both figures are client-side estimates and can differ from your actual bill.

Piped stdin is capped at 10MB. If you exceed the cap, Claude Code exits with a clear error and a non-zero status. To work with larger inputs, write the content to a file and reference the file path in your prompt instead of piping it.

If Claude Code can't read stdin, for example because the process that started it disconnected its end, Claude Code prints a warning to stderr and continues with the prompt from the command line. Before v2.1.211, an unreadable stdin on Windows crashed the session or made it exit silently with no output.

Add Claude to a build script

You can wrap a non-interactive call in a script to use Claude as a project-specific linter or reviewer.

This package.json script pipes the diff against main into Claude and asks it to report typos. Piping the diff means Claude doesn't need Bash permission to read it, and the escaped double quotes keep the script portable to Windows:

{
  "scripts": {
    "lint:claude": "git diff main | claude -p \"you are a typo linter. for each typo in this diff, report filename:line on one line and the issue on the next. return nothing else.\""
  }
}

Run it with npm run lint:claude.

Get structured output

Use --output-format to control how responses are returned:

  • text (default): plain text output
  • json: structured JSON with result, session ID, and metadata
  • stream-json: newline-delimited JSON for real-time streaming

This example returns a project summary as JSON with session metadata, with the text result in the result field:

claude -p "Summarize this project" --output-format json

To get output conforming to a specific schema, use --output-format json with --json-schema and a JSON Schema definition. The response includes metadata about the request (session ID, usage, etc.) with the structured output in the structured_output field.

This example extracts function names and returns them as an array of strings:

claude -p "Extract the main function names from auth.py" \
  --output-format json \
  --json-schema '{"type":"object","properties":{"functions":{"type":"array","items":{"type":"string"}}},"required":["functions"]}'

If the value isn't a valid JSON Schema, claude exits with Error: --json-schema is not a valid JSON Schema followed by the validator's diagnostic. Claude Code accepts schemas that use the format keyword, such as "format": "email", but treats format as an annotation and doesn't enforce it. Before v2.1.205, Claude Code silently ignored an invalid schema and returned unstructured text, and treated any schema containing format as invalid.

Use a tool like [jq](https://jqlang.org/) to parse the response and extract specific fields:
# Extract the text result
claude -p "Summarize this project" --output-format json | jq -r '.result'

# Extract structured output
claude -p "Extract function names from auth.py" \
  --output-format json \
  --json-schema '{"type":"object","properties":{"functions":{"type":"array","items":{"type":"string"}}},"required":["functions"]}' \
  | jq '.structured_output'

Stream responses

Use --output-format stream-json with --verbose and --include-partial-messages to receive tokens as they're generated. Each line is a JSON object representing an event:

claude -p "Explain recursion" --output-format stream-json --verbose --include-partial-messages

The last line of the stream is a result message with the final response text, cost, and session metadata.

If your consumer reads the stream slowly, Claude Code waits for the queued output to drain before exiting, scaling the wait with how much is still queued, capped at 30 seconds. Before v2.1.214 the exit wait was capped at about two seconds, which could cut off the end of a large response.

The following example uses jq to filter for text deltas and display just the streaming text. The -r flag outputs raw strings (no quotes) and -j joins without newlines so tokens stream continuously:

claude -p "Write a poem" --output-format stream-json --verbose --include-partial-messages | \
  jq -rj 'select(.type == "stream_event" and .event.delta.type? == "text_delta") | .event.delta.text'

For programmatic streaming with callbacks and message objects, see Stream responses in real-time in the Agent SDK documentation.

Follow subagent messages

Messages from subagents appear in the stream as assistant and user messages whose parent_tool_use_id field is the ID of the tool call that spawned the subagent. Messages from the main conversation carry null in that field.

By default, Claude Code emits only subagent tool_use and tool_result blocks. Pass --forward-subagent-text or set CLAUDE_CODE_FORWARD_SUBAGENT_TEXT to also emit subagent text and thinking blocks, so you can reconstruct each subagent's transcript. This requires Claude Code v2.1.211 or later.

When you enable either option, Claude Code forwards messages from subagents at every nesting depth: when a subagent spawns its own subagent, the nested s …[truncated 16635 chars]

[FETCH #26] https://urlwatch.readthedocs.io/en/latest/jobs.html

ASK: Extract: jobs.yaml format for url jobs and shell jobs, key options (url, name, filter, headers, method, data, cookies, encoding, ignore_connection_errors, use_browser), and example snippets. Also note where config is stored on Windows and how to run urlwatch.

urlwatch Jobs Configuration Guide

URL Jobs Format

URL jobs monitor web documents. Basic structure:

name: "urlwatch homepage"
url: "https://thp.io/2008/urlwatch/"

Key options for URL jobs:

  • url (required): Target document address
  • name: Human-readable identifier
  • method: HTTP verb (default: GET)
  • data: POST/PUT payload
  • headers: Custom HTTP headers
  • cookies: Request cookies
  • encoding: Character encoding override
  • filter: Data transformation rules
  • ignore_connection_errors: Suppress connection failures

Shell Jobs Format

Monitor arbitrary command output:

name: "What is in my Home Directory?"
command: "ls -al ~"

Shell jobs execute system commands and track output changes, useful for monitoring folders or device status.

Configuration Storage & Execution

Jobs reside in urls.yaml, accessed via urlwatch --edit. The urlwatch --list command displays all configured jobs with assigned index numbers. Each job entry is separated by a line containing only ---.

Additional Configuration

The main configuration file supports a job_defaults section to apply settings across all jobs simultaneously, eliminating repetitive key declarations.

[FETCH #27] https://changedetection.io/docs/api_v1/index.html

ASK: Extract: API endpoints for creating a watch, listing watches, fetching watch history/snapshot, triggering recheck, notifications, and how the x-api-key header is used. Provide example curl.

ChangeDetection.io API Overview

Core Authentication

All API requests require an API key passed via the x-api-key header. Retrieve your key from Settings > API in the dashboard.

Key Endpoints

Watch Management

List all watches:

curl -X GET "http://localhost:5000/api/v1/watch" \
  -H "x-api-key: YOUR_API_KEY"

Create a new watch:

curl -X POST "http://localhost:5000/api/v1/watch" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "title": "My Monitor",
    "time_between_check": {"hours": 1}
  }'

Get single watch:

curl -X GET "http://localhost:5000/api/v1/watch/{uuid}" \
  -H "x-api-key: YOUR_API_KEY"

Update watch:

curl -X PUT "http://localhost:5000/api/v1/watch/{uuid}" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"notification_muted": true}'

Delete watch:

curl -X DELETE "http://localhost:5000/api/v1/watch/{uuid}" \
  -H "x-api-key: YOUR_API_KEY"

Monitoring Data

Fetch watch history:

curl -X GET "http://localhost:5000/api/v1/watch/{uuid}/history" \
  -H "x-api-key: YOUR_API_KEY"

Get snapshot:

curl -X GET "http://localhost:5000/api/v1/watch/{uuid}/history/latest" \
  -H "x-api-key: YOUR_API_KEY"

Compare snapshots:

curl -X GET "http://localhost:5000/api/v1/watch/{uuid}/difference/previous/latest?format=htmlcolor" \
  -H "x-api-key: YOUR_API_KEY"

Triggering Actions

Force recheck: Add ?recheck=1 to the GET watch endpoint.

Toggle pause/mute: Use query parameters ?paused=paused or ?muted=muted.

Notifications

The API supports per-watch notification URLs using "apprise" format (email, Discord, Slack, webhooks, etc.).

Global notification management:

curl -X POST "http://localhost:5000/api/v1/notifications" \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"notification_urls": ["mailto:admin@example.com"]}'

Connection URL Format

  • Local: http://localhost:5000/api/v1/
  • Hosted: https://<your-domain>/api/v1/

All operations follow standard HTTP methods (GET, POST, PUT, DELETE) with JSON payloads.

[FETCH #28] https://nedrug.mfds.go.kr/bbs/117

ASK: Describe this page: what board is it (DMF 원료의약품등록 정보), list columns, pagination pattern, any search form fields, and post titles seen (e.g. 등록 공고, 변경, 취하).

Page Description: DMF 원료의약품등록 정보 Board

Board Type: Drug Substance Registration (DMF) Information announcement board

Table Columns:

  • 연번 (Serial Number)
  • 제목 (Title)
  • 조회건수 (View Count)
  • 등록자 (Registrant)
  • 등록일자 (Registration Date)

Pagination Pattern: First page | Previous | Pages 1-10 | Next | Last page (total 710 posts)

Display Options: 10, 20, 30, 40, or 50 posts per page

Post Title Pattern: Entries follow a consistent format: "등록대상 원료의약품(DMF) 등록 공고" (Drug Substance Registration Announcement) with specific date ranges, organized weekly. Sample titles shown range from February 2021 back to November 2020.

Search Form: The page includes a "제목" (Title) search field with search and reset buttons at the top.

No advanced filters or specific modification/withdrawal post types are visible in the displayed results.

[FETCH #29] https://www.data.go.kr/data/15057075/openapi.do

ASK: Extract: API name, endpoint URL, operation names, request parameters (serviceKey, pageNo, numOfRows, type, etc.), response fields (등록번호, 성분명, 업체명, 제조소명, 발급일자 etc.), update cycle, and sample code if present.

API Extraction: DMF Registration Status

API Details

Name: 식품의약품안전처_원료의약품등록(DMF)현황 (Food and Drug Safety DMF Registration Status)

Endpoint URL: https://apis.data.go.kr/1471000/MdcDmfInfoService01/getMdcDmfList01

Service URL: https://apis.data.go.kr/1471000/MdcDmfInfoService01

Request Parameters

Parameter Type Required Description
serviceKey string Yes Authentication key from the data portal
pageNo integer No Page number (default: 1)
numOfRows integer No Results per page (default: 3)
entp_name string No Company/manufacturer name
ingr_kor_name string No Ingredient name (Korean)
type string No Response format: xml or json

Response Fields

  • DMF_PERMIT_NO Registration number
  • INGR_KOR_NAME Ingredient name
  • ENTP_NAME Company name
  • MNFCTR_NAME Manufacturing facility name
  • MNFCTR_PLACE Manufacturing facility location
  • MANUF_COUNTRY_CODE_NM Manufacturing country
  • DMF_PERMIT_DATE Issue date
  • resultCode Status code
  • resultMsg Status message
  • totalCount Total result count

Additional Information

  • Format: XML/JSON
  • Cost: Free
  • Traffic Limit (Dev): 10,000 calls
  • Update Frequency: Last modified September 19, 2025
  • Provider: Ministry of Food and Drug Safety

[FETCH #30] https://github.com/mrueda/nomenclator-delta

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (which data source, how monthly delta/diff is computed: added/removed/changed, output format, scheduling, CLI usage). Quote numbers precisely.

nomenclator-delta Repository Summary

Repository Description: "Compare monthly changes in medicines and health products from the Ministerio de Sanidad's Nomenclátor de Facturación."

Metrics:

  • Stars: 1
  • Forks: 0
  • Primary Language: Python

Last Commit: Not explicitly stated in provided content

File Structure:

src/nomenclator_delta/     (collection, normalization, diffing, validation)
data/                      (snapshots and change history)
site/                      (Spanish static application)
docs-site/                 (documentation and Pages build)
tests/                     (unit and integration tests)

Data Source: The application monitors Spain's "Nomenclátor de facturación" from the Ministry of Health's public database of medicines in the National Health System.

Delta/Diff Computation: The tool "compares consecutive monthly releases" and "finds changes by Código Nacional, medicine name, active ingredient, or laboratory." The README indicates capabilities for identifying added, removed, and changed entries, though specific algorithmic details aren't elaborated in the excerpt provided.

Output Format: A static browser application with no backend; users can "Open the app" at the hosted GitHub Pages site.

Scheduling/Updates: Documentation references "a monthly update runbook" for maintainers managing routine releases.

CLI Usage:

python3 -m nomenclator_delta validate data
python3 -m nomenclator_delta dist

License: MIT

[FETCH #31] https://github.com/Mzands2622/Zanalytix

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (how scraping works, how LLM-based change detection compares old vs new snapshots, storage, scheduling, notification). Quote numbers precisely.

Zanalytix Repository Summary

Repo Description: "AI-powered pipeline for scraping and tracking pharmaceutical clinical pipeline data with LLM-based change detection."

Stats:

  • Stars: 0
  • Forks: 0
  • Primary Language: Python
  • Last Commit: 1 commit on main branch (specific date not provided)

File Structure

The repository contains 60+ company-specific parser modules (e.g., abbvie_pipeline.py, pfizer_pipeline.py) alongside core infrastructure files:

  • Database management (db.py)
  • API layer (function_app.py, master_scheduler.py)
  • User authentication (login.py, sign_up.py)
  • Notification systems (notifications.py)
  • Utility scripts (cleanup_text.py, treatment_visualizer.py)

How Scraping Works

"HTML Scraping (Zyte API) --> 60+ Company Parsers (BeautifulSoup)" feeds into a standardization pipeline. Each company module implements two functions: fetch_{company}_html() retrieves pipeline pages, while process_{company}_html() parses results into structured objects.

LLM-Based Change Detection

"GPT-4o compares old vs new treatment data, assigns priority (1-5), categorizes changes" via the comparison queue. The system maintains dated snapshots in JSON format, enabling temporal analysis.

Storage Architecture

The Revised_MasterTable stores treatments as keys with Treatment_Data containing timestamped snapshots. A separate Stream table tracks GPT-4o responses and metadata per change event.

Scheduling & Notifications

"Calendar-based scraping schedules with recurrence rules and auto-extension" trigger Azure Functions. Matched changes route to users via "Twilio SMS/Calls, Email" based on saved preferences.

[FETCH #32] https://github.com/larsyencken/csvdiff

ASK: Extract: repo description, star count, fork count, primary language, last commit date, CLI usage (csvdiff KEY a.csv b.csv, --style, JSON output with added/removed/changed), Python API usage. Quote numbers precisely.

csvdiff Repository Summary

Description: "Generate a diff between two tabular datasets expressed in CSV files."

Stars: 131

Forks: 31

Primary Language: Python

Status: Archived February 18, 2021 (read-only, no longer maintained)

CLI Usage

Basic syntax:

csvdiff --style=summary KEY file1.csv file2.csv

Key options:

  • --style: Choose output format (summary, pretty)
  • --output: Specify JSON output file
  • --ignore-columns: Comma-separated list of columns to exclude
  • --significance: Numeric precision for comparison (negative values for orders of magnitude)

Example:

csvdiff --style=summary id a.csv b.csv

JSON Output Structure

The diff output contains:

  • _index: Array of key columns
  • added: New rows with all fields
  • removed: Deleted rows with all fields
  • changed: Modified rows showing field-level "from/to" changes organized by key

Python API

Main entry points:

import csvdiff
patch = csvdiff.diff_files('a.csv', 'b.csv', ['id'])
patch = csvdiff.diff_records(records_a, records_b, ['id'])

Matching patch methods available for applying diffs to data.

License: BSD-3-Clause

[FETCH #33] https://github.com/Windos/BurntToast

ASK: Extract: repo description, star count, fork count, primary language, last release/version, install command (Install-Module), basic usage (New-BurntToastNotification -Text ..., buttons via New-BTButton, -AppLogo), any note about running from SYSTEM/service context. Quote numbers precisely.

BurntToast Repository Summary

Description: "PowerShell Module for displaying Toast Notifications on Windows 10 and Windows Server 2019 and above."

Metrics:

  • Stars: 1.7k
  • Forks: 126
  • Primary Language: PowerShell

Latest Release: v1.1.0

  • Notable additions include support for urgent notifications via the Urgent switch and button color customization

Installation:

Install-Module -Name BurntToast

Basic Usage: The module provides cmdlets like New-BurntToastNotification for creating notifications and New-BTButton for adding interactive buttons. Users can specify app branding through the -AppLogo parameter.

Important Limitation: The module targets user-context notifications and has limitations when running from SYSTEM or service accounts due to the Windows notification framework's desktop session requirements.

License: MIT

[FETCH #34] https://github.com/michalzobec/autorunsalerts

ASK: Extract: repo description, star count, primary language, last commit date, how the two scheduled tasks (SYSTEM scan + user-context toast) are structured, files. Quote numbers precisely.

AutorunsAlerts Repository Summary

Description: "Simple toast notifications for changes to autoruns configurations on windows"

Stars: 0

Primary Language: PowerShell

Last Commit Date: Not explicitly stated in the provided content

Scheduled Tasks Architecture

The system uses two complementary tasks:

  1. AutorunsAlert (SYSTEM context): Executes every 60 minutes as the system account. It runs autorunsc.exe to scan the current autoruns state, compares findings against the previous baseline stored in state.json, and logs all detected changes to audit.log.

  2. AutorunsAlertToast (User context): Runs every 15 minutes within the user session. This task checks for notification flags set by the scanner and generates toast alerts that link to audit.log for user investigation.

Key Files

  • autorunsalert.ps1 — Main scanning script
  • autorunstoast.ps1 — Toast notification delivery script
  • configuration.json — Shared configuration variables
  • state.json — Previous scan baseline for comparison
  • audit.log — Change history and investigation records
  • install.ps1 and uninstall.ps1 — Setup/teardown scripts

This separation ensures alerts appear in the user's active session rather than the system session where detection occurs.

[FETCH #35] https://github.com/addyosmani/gemini-cli-tips

ASK: Extract: repo description, star count, fork count, last commit date, and the exact section/snippet about headless mode (-p), cron scheduling, --output-format json, and any pipeline examples. Quote exact commands.

Gemini CLI Tips Repository Analysis

Repository Metadata

  • Description: Gemini CLI Tips and Tricks
  • Star Count: 2.4k
  • Fork Count: 105
  • Last Commit Date: Not explicitly provided in the content

Headless Mode Details

The guide covers running Gemini CLI non-interactively using the -p flag:

One-shot execution:

gemini -p "Your prompt here"

This approach allows you to "output a single response and exit" without entering the interactive REPL.

Input piping: The content demonstrates Unix-style stdin integration: echo "Count to 10" | gemini feeds prompts via pipe.

JSON Output & Automation

The guide mentions using --format=json for programmatic consumption, enabling scripts to "parse the JSON to get the answer or any tool actions details."

Pipeline & Scheduling Context

While specific cron examples aren't provided, the guide suggests headless mode enables broader automation: "It transforms Gemini CLI from an interactive assistant into a backend service or utility that other programs can call."

The documentation implies you could "schedule a cronjob that runs a Gemini CLI prompt nightly" for automated tasks like report generation or system cleanup.

System Prompt Override

Advanced users can customize behavior via environment variable:

export GEMINI_SYSTEM_MD="/path/to/custom_system.md"

This replaces the default system instructions, enabling specialized workflows or persona enforcement.

[FETCH #36] https://github.com/NomaDamas/k-skill/blob/main/docs/features/mfds-food-safety.md

ASK: Extract: what this doc describes (which MFDS APIs, scripts, how API keys are handled, python scripts paths), and the parent repo purpose. Also note stars if visible.

Food Safety Check Guide Documentation

Repository Overview

The k-skill repository (7.4k stars) is a GitHub project by NomaDamas that provides skill modules with helper scripts and CLI tools.

Document Purpose

This markdown file documents the "식품 안전 체크 가이드" (Food Safety Check Guide), a feature that enables users to verify food products against official Korean food safety databases.

MFDS APIs Referenced

The document references two primary data sources:

  1. 공공데이터포털 (Public Data Portal): "부적합 식품" (unsuitable food) endpoint for retrieving non-compliant products
  2. 식품안전나라 (Food Safety Korea): Recall and sales suspension API for tracking withdrawn products

Script & API Key Handling

Script Location: scripts/mfds_food_safety.py (included with the skill installation)

API Key Architecture: Credentials (DATA_GO_KR_API_KEY and FOODSAFETYKOREA_API_KEY) must be configured on the "프록시 운영 서버" (proxy server) environment, not on the user's side. Users access functionality through k-skill-proxy at /v1/mfds/food-safety/search.

Key Implementation Detail

The guide emphasizes: "이 helper는 직접 진단을 하지 않는다" (this helper does not provide direct diagnosis), prioritizing medical consultation over automated assessment when symptoms are reported.

[FETCH #37] https://github.com/WooilJeong/PublicDataReader

ASK: Extract: repo description, star count, fork count, primary language, last commit date, install command, list of supported public data providers/APIs (especially any 식약처/MFDS/의약품 coverage), usage example. Quote numbers precisely.

PublicDataReader Repository Summary

Repository Description: "공공 데이터 조회를 위한 오픈소스 파이썬 라이브러리" (Open-source Python library for retrieving public data)

Repository Metrics:

  • Stars: 597
  • Forks: 113
  • Primary Language: Python

Installation Command:

pip install PublicDataReader --upgrade

Supported Public Data Providers:

The library provides access to the following APIs:

  • FRED (Federal Reserve Economic Data)
  • 공공데이터포털 (Public Data Portal) - multiple datasets including:
    • 국토교통부 (Ministry of Land, Infrastructure & Transport): Real estate transaction prices, building ledgers, construction permits, housing permits, land/forest records, land ownership
    • 소상공인시장진흥공단 (Small Business & Market Administration): Commercial district information
    • 한국자산관리공사 (KAMCO): Public auction properties
    • 국세청 (NTS): Business registration verification
    • 한국부동산원 (Korea Real Estate Board): Real estate information
  • KOSIS (국가통계포털 - National Statistics Portal)
  • ECOS (한국은행 - Bank of Korea Economic Statistics)
  • 서울시 (Seoul City): Transportation data
  • V-World (국가공간정보 - National Spatial Information)
  • KB부동산 (KB Real Estate): Housing price trends

Note: No 식약처/MFDS (Korean FDA) or pharmaceutical coverage is mentioned in the documentation.

Last Commit: 168 commits recorded; specific date not provided in page content

[FETCH #38] https://github.com/caronc/apprise

ASK: Extract: repo description, star count, fork count, primary language, last release, install command, Python usage example (apprise.Apprise(), add(), notify()), and whether Windows toast / Slack / Telegram / email are supported with URL formats. Quote numbers precisely.

Apprise Repository Summary

Description: "Push Notifications that work with just about every platform!"

Stars: 17.2k | Forks: 652

Primary Language: Python

Installation:

pip install apprise

Python Usage Example:

import apprise
apobj = apprise.Apprise()
apobj.add('mailto://myuserid:mypass@gmail.com')
apobj.notify(body='notification text', title='my title')

Supported Services:

  • Windows Toast: Yes - windows://
  • Slack: Yes - slack://TokenA/TokenB/TokenC/Channel
  • Telegram: Yes - tgram://bottoken/ChatID
  • Email: Yes - mailto://userid:pass@domain.com and mailtos:// (secure)

Last Release Information: Not explicitly stated in the provided content, though the repository shows "1,178 Commits" on the master branch with active development indicated by recent GitHub Actions workflows.

The library supports over 100+ notification services through a unified URL-based syntax, enabling developers to integrate multiple notification channels without service-specific implementations.

[FETCH #39] https://github.com/jshchnz/claude-code-scheduler/tree/main/examples

ASK: List the example files and summarize the JSON schedule format shown, especially any Windows Task Scheduler-specific example, cron expression, execution.command, timeout, skipPermissions fields.

Claude Code Scheduler Examples Overview

Based on the GitHub page content provided, the examples directory contains two JSON schedule files:

  1. daily-review.json
  2. weekly-audit.json

Limitations of Available Information

Unfortunately, the webpage content shown does not include the actual JSON file contents or details about the schedule format structure. The page displays only the file listing in the GitHub repository directory view, without revealing:

  • The specific JSON schema or fields used
  • Windows Task Scheduler-specific configuration examples
  • Cron expression formats
  • Details about execution.command, timeout, or skipPermissions fields
  • How scheduling intervals are defined

To examine the actual format and field specifications, you would need to view the raw JSON files directly in the repository by clicking on the individual files (daily-review.json or weekly-audit.json) on GitHub.

[FETCH #40] https://gist.github.com/drmalex07/10554232

ASK: Extract the full pywin32 service example code structure: class inheriting win32serviceutil.ServiceFramework, svc_name, svc_display_name, SvcStop, SvcDoRun, main loop with win32event.WaitForSingleObject, and the main HandleCommandLine call. Quote code.

PyWin32 Windows Service Example Structure

Based on the GitHub Gist, here's the core service architecture:

Class Definition

The example implements a service by subclassing the framework:

class HelloWorldSvc (win32serviceutil.ServiceFramework):
    _svc_name_ = "HelloWorld-Service"
    _svc_display_name_ = "HelloWorld Service"

Initialization

def __init__(self,args):
    win32serviceutil.ServiceFramework.__init__(self,args)
    self.stop_event = win32event.CreateEvent(None,0,0,None)
    socket.setdefaulttimeout(60)
    self.stop_requested = False

Key Methods

The SvcStop method handles shutdown signals: "self.ReportServiceStatus(win32service.SERVICE_STOP_PENDING)" and sets the stop flag.

The SvcDoRun method launches the service's operational logic, calling self.main().

Main Loop

The service runs an iterative loop checking if self.stop_requested: to break execution when signaled.

Entry Point

if __name__ == '__main__':
    win32serviceutil.HandleCommandLine(HelloWorldSvc)

Note: Comments in the thread suggest this basic pattern may encounter timeout errors; enhanced implementations check len(sys.argv) to differentiate between command-line installation versus actual service execution.

[FETCH #41] https://github.com/HasData/playwright-scraping

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure (python/ and nodejs/ scripts), what scraping tasks are demonstrated. Quote numbers precisely.

Playwright Scraping Repository Summary

Repository Description: "Web scraping and browser automation using Playwright in both Python and Node.js. It includes scripts for common tasks such as scraping data, interacting with web elements, handling authentication, and managing errors."

Engagement Metrics:

  • Stars: 15
  • Forks: 4

Primary Languages: Python and Node.js (dual implementation)

Last Commit: The repository shows "5 Commits" on the main branch, though a specific date is not visible in the provided content.

Directory Structure: Both Python/ and NodeJS/ folders contain identical organizational patterns:

  • basics/ (browser launching, headless modes, multiple tabs)
  • scraping/ (text, links, images, Shadow DOM, waiting)
  • selectors/ (CSS, XPath, role-based, text-based)
  • interactions/ (clicking, forms, dropdowns, pagination, scrolling)
  • save_data/ (JSON, CSV, PDF, downloads, screenshots)
  • auth/ (basic authentication, cookies)
  • browser/ (user agents, proxies, device emulation)
  • errors/ (retry logic, timeout handling)
  • debug/ (video/trace recording, pausing, console inspection)

Demonstrated Scraping Tasks: The repository covers extraction of product data from Amazon and WooCommerce platforms, element selection techniques, form interactions, pagination handling, and infinite scroll implementation.

[FETCH #42] https://github.com/Tanguy9862/AI-Powered-FDA-Drug-Scraper

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (which FDA page scraped, how LangChain/GPT classify, output format). Quote numbers precisely.

AI-Powered FDA Drug Scraper - Repository Analysis

Repository Metadata

  • Description: "Python-based web scraper leveraging generative AI with LangChain and GPT-4o-mini to extract and classify FDA drug approval data"
  • Stars: 3
  • Forks: 0
  • Primary Language: Python
  • Last Commit Date: Not explicitly stated in provided content

File Structure

The repository contains:

  • new_drug_approvals_scraper/ (main package directory)
  • img_readme/ (documentation images)
  • scraper.py, classification.py, utils.py, __init__.py
  • requirements.txt, setup.py, .gitignore, LICENSE, README.md

Data Source

The scraper targets "Drugs.com" for extraction of new drug approval information.

Classification Methodology

LangChain integrates with OpenAI's GPT models to "categoriz[e] medications and their treatment categories." The system processes "over 1,770 records" and applies "logical rules to categorize drugs based on their detailed descriptions and intended uses."

Output & Processing

The scraper performs data normalization reducing company name variations from "approximately 1000 to 700" unique entries through standardization of suffixes, abbreviations, and collaboration descriptors.

Integration

A companion project exists: the "New Drug Approvals Dashboard repository" demonstrates real-time implementation within Dash applications.

[FETCH #43] https://github.com/ecprice/newsdiffs

ASK: Extract: repo description, star count, fork count, primary language, last commit date, architecture (how snapshots are stored: git repo per article? diff generation; scheduling by cron). Quote numbers precisely.

NewsDiffs Repository Analysis

Repo Description: "A website and framework that tracks changes in online news articles over time."

Star Count: 506 stars

Fork Count: 136 forks

Primary Language: Python (evidenced by .py files throughout and Django framework usage)

Last Commit Date: Not explicitly stated in the provided content

Architecture:

The system uses a directory-based storage approach rather than per-article git repositories. Key components include:

  • Snapshot Storage: Articles are stored in an articles directory created during setup
  • Diff Generation: The scraper runs periodically to capture article versions and detect changes between snapshots
  • Scheduling: Uses either manual cron jobs or loop-based scheduling: "while true; do python website/manage.py scraper; sleep 60m; done" for hourly execution
  • Parser Framework: Modular parser architecture in parsers/ directory, with site-specific subclasses extending BaseParser
  • Logging: Progress logged to /tmp/newsdiffs_logging (per-run); errors to /tmp/newsdiffs/logging_errs (cumulative)

The system compares successive scrapes to identify and display article modifications over time.

[SEARCH #27] github FDA "dmflist" xlsx drug master file list download python quarterly

Web search results for query: "github FDA "dmflist" xlsx drug master file list download python quarterly"

Links: [{"title":"List of Drug Master Files (DMFs) | FDA","url":"https://www.fda.gov/drugs/drug-master-files-dmfs/list-drug-master-files-dmfs"},{"title":"FDA Drug Master Files Directory - John Snow Labs","url":"https://www.johnsnowlabs.com/marketplace/fda-drug-master-files-directory/"},{"title":"Drug Master File","url":"https://en.wikipedia.org/wiki/Drug_Master_File"},{"title":"US Drug Master Files (DMF) Database | PharmaCompass.com","url":"https://www.pharmacompass.com/us-drug-master-files-dmfs"},{"title":"Drug Master Files (DMFs) | FDA","url":"https://www.fda.gov/drugs/forms-submission-requirements/drug-master-files-dmfs"},{"title":"GitHub - vshah1016/pharma_scraper: A scraper I made in python to take stocks that are going to have a drug released and have all of its data aggregated to a CSV file where I can sort and make big bux","url":"https://github.com/vshah1016/pharma_scraper"},{"title":"Excipient DMF List - pharma excipients","url":"https://www.pharmaexcipients.com/excipient-sources/excipient-dmf-list/"},{"title":"Dmf List FDA Quarterly Drug Master Files Spreadsheet","url":"https://dmf-list.backgroundscheck.info/"},{"title":"Drug Master File - US DMF FDA Guidance and Submission","url":"https://fdapals.com/services/dmf-fda-guidance/"}]

Based on the search results, here's what I found:

The FDA maintains a list of DMFs, which is updated quarterly, containing DMFs received by June 30, 2026, for which acknowledgment letters were sent before July 19, 2026.

The DMF List provides a quarterly FDA spreadsheet of Drug Master Files, including DMF numbers, submitters, file type, acknowledgment date and review division. The spreadsheet is freely available on the FDA's DMF webpage and is refreshed quarterly to add newly acknowledged files or note withdrawals.

The US FDA maintains a Public Database of all Drug Master Files filed with US FDA, anyone can access this database which will download an excel file, and the USFDA updates the list every quarter.

The search results show that the FDA provides official quarterly DMF lists available as downloadable spreadsheets on their website. However, the results don't show a specific Python library or GitHub repository specifically named "dmflist" for downloading this data. The search results reference general pharmaceutical scraping tools on GitHub but not a specific tool dedicated to automating DMF list downloads.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #28] schtasks OR "Task Scheduler" "claude -p" windows powershell scheduled task Claude Code automation

Web search results for query: "schtasks OR "Task Scheduler" "claude -p" windows powershell scheduled task Claude Code automation"

Links: [{"title":"GitHub - jshchnz/claude-code-scheduler: Put Claude on autopilot · GitHub","url":"https://github.com/jshchnz/claude-code-scheduler"},{"title":"Schedule recurring tasks in Claude Code Desktop - Claude Code Docs","url":"https://code.claude.com/docs/en/desktop-scheduled-tasks"},{"title":"Windows Task Scheduler Claude Code Skill | AI Automation","url":"https://mcpmarket.com/tools/skills/windows-task-scheduler"},{"title":"Claude Code Scheduled Tasks: Complete Setup Guide (2026)","url":"https://claudefa.st/blog/guide/development/scheduled-tasks"},{"title":"Scheduled Tasks: How to Put Claude on Autopilot | atal upadhyay","url":"https://atalupadhyay.wordpress.com/2026/03/02/scheduled-tasks-how-to-put-claude-on-autopilot/"},{"title":"Mastering Scheduled Tasks in Claude Code: A Complete Guide | aixplore","url":"https://aixplore.in/blog_post?slug=mastering-scheduled-tasks-in-claude-code-guide"},{"title":"Claude Cowork Scheduled Tasks: Automate Your Recurring Workflows - Claude Cowork Blog | Claude Cowork Guide","url":"https://claudecowork.im/blog/scheduled-tasks-guide"}]

Based on the search results, here's what I found about using schtasks and Task Scheduler with Claude Code:

Windows Task Scheduler Integration

At the scheduled time, tasks run claude -p "your command" (with --dangerously-skip-permissions for autonomous tasks), and on Windows, you can check scheduled Claude tasks using schtasks /query /tn "ClaudeSchedule*".

Claude Code Scheduler Plugin

Claude Code supports cross-platform scheduling including Windows Task Scheduler (in addition to macOS launchd and Linux crontab). The Claude Code scheduler enables you to schedule code reviews, security audits, and other tasks that run automatically - the scheduler works on macOS, Linux, and Windows.

How It Works

When you create a task using the plugin, Claude generates a shell script (called a "wrapper") that calls Claude Code with your prompt as input, then registers this wrapper with Task Scheduler at the interval you specify. Since Task Scheduler is a system-level process, it runs independently of any application — your task executes even if you never open the Claude app.

Windows Task Scheduler Skill

There's also a Task Scheduler skill that enables creating and managing Windows Task Scheduler jobs using Claude Code, automating reminders, recurring scripts, and system tasks with PowerShell integration.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #29] github python windows Task Scheduler scraper "schtasks" create daily xlsx report script repository

Web search results for query: "github python windows Task Scheduler scraper "schtasks" create daily xlsx report script repository"

Links: [{"title":"Python script to interact with existing Windows Task Scheduler tasks. · GitHub","url":"https://gist.github.com/nmpowell/dc8e7187948788c5c126f01755252164"},{"title":"task-scheduler · GitHub Topics · GitHub","url":"https://github.com/topics/task-scheduler?l=powershell"},{"title":"GitHub - 786raees/task-scheduler-python: The Task Scheduler Python project provides a convenient way to interact with the Windows Task Scheduler using Python and the win32com.client library. You can easily create, manage, toggle, run, and delete scheduled tasks, making task automation a breeze. Explore the methods provided by the TaskScheduler class to customize. · GitHub","url":"https://github.com/786raees/task-scheduler-python"},{"title":"Two ways to run Python Scripts Every Day Automatically | by Vnalla | Medium","url":"https://medium.com/@vineelan09/two-ways-to-run-python-scripts-every-day-automatically-3c86079fe449"},{"title":"Automated Web Scraper With Python & Windows Task Scheduler","url":"https://oxylabs.io/blog/automated-web-scraper-windows-task-scheduler"},{"title":"Run a python script on a schedule using the built-in Task Scheduler Windows app | Biztory","url":"https://biztory.com/blog/run-a-python-script-on-a-schedule-using-the-in-built-task-scheduler-windows-app"},{"title":"How to Automate Python Scripts with Task Scheduler (Windows example) - JC Chouinard","url":"https://www.jcchouinard.com/python-automation-using-task-scheduler/"},{"title":"Scheduling | Adobe User Sync Tool","url":"https://adobe-apiplatform.github.io/user-sync.py/en/success-guide/scheduling.html"},{"title":"Automated Web Scraper With Python & Windows Task Scheduler - Flipnode","url":"https://flipnode.io/automated-web-scraper-windows-task-scheduler"},{"title":"Making Your Python Web Scraper Work for You: Automating With Windows Task Scheduler - Oreate AI Blog","url":"https://www.oreateai.com/blog/making-your-python-web-scraper-work-for-you-automating-with-windows-task-scheduler/7efa97048b3b0b28a3513e73ed1235fd"}]

I found several relevant search results related to Python scripts for Windows Task Scheduler automation. Here are the key findings:

GitHub Repositories & Projects:

  1. Task Scheduler Python Project - There's a Task Scheduler Python project that provides a convenient way to interact with the Windows Task Scheduler using Python and the win32com.client library, allowing you to easily create, manage, toggle, run, and delete scheduled tasks. This is available at: https://github.com/786raees/task-scheduler-python

  2. Windows Task Scheduler Python Script - A Python script to interact with existing Windows Task Scheduler tasks is available as a GitHub Gist at: https://gist.github.com/nmpowell/dc8e7187948788c5c126f01755252164

Using schtasks Command:

The schtasks command can be used to create daily tasks with syntax like: schtasks /create /tn "Task Name" /tr path_to_bat_file/run.bat /sc DAILY /st 16:00

Creating Daily XLSX Reports:

There are examples of ways to log system events to both txt and xlsx files on Windows machines, and scripts can be used to run reporting scripts in the proper order and on the correct date.

Practical Implementation:

A common approach is to create a Batch file using a text editor that invokes Python and runs the script, then schedule it using Windows Task Scheduler to run at regular intervals like daily.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[SEARCH #30] github 공공데이터포털 data.go.kr 파이썬 라이브러리 PublicDataReader 식약처 의약품

Web search results for query: "github 공공데이터포털 data.go.kr 파이썬 라이브러리 PublicDataReader 식약처 의약품"

Links: [{"title":"식품의약품안전처 식품의약품안전평가원_식품의약품안전처 연구관리 기술 분류 정보조회 서비스 | 공공데이터포털","url":"https://www.data.go.kr/data/15068423/openapi.do"},{"title":"GitHub - WooilJeong/PublicDataReader: 공공 데이터 조회를 위한 오픈소스 파이썬 라이브러리 · GitHub","url":"https://github.com/WooilJeong/PublicDataReader"},{"title":"GitHub - jjscan/data.go.kr-1: 공공데이터(www.data.go.kr) - 식품의약품안전처 의약품 관련 정보-DUR품목정보(DURPrdlstInfoService)-병용금기정보조회(getUsjntTabooInfoList) 데이터 수집","url":"https://github.com/jjscan/data.go.kr-1"},{"title":"식품의약품안전처 식품의약품안전평가원_식품의약품안전처 연구관리 전문기술분야코드 조회 서비스 | 공공데이터포털","url":"https://www.data.go.kr/data/15068280/openapi.do"},{"title":"GitHub - Q00/data.go.kr-crawling: 건강정보, 의약품 크롤링, 멀티쓰레딩 gevent · GitHub","url":"https://github.com/Q00/data.go.kr-crawling"},{"title":"식의약 데이터 포털","url":"https://data.mfds.go.kr/"},{"title":"식품의약품안전처 의약품 관련 정보_20201031 | 공공데이터포털","url":"https://www.data.go.kr/data/15020627/fileData.do"},{"title":"식품의약품안전처_의약품 제품 허가정보 | 공공데이터포털","url":"https://www.data.go.kr/data/15095677/openapi.do"},{"title":"식품의약품안전처_의약품 낱알식별 정보 | 공공데이터포털","url":"https://www.data.go.kr/data/15057639/openapi.do?recommendDataYn=Y"},{"title":"식품의약품안전처_필수의약품내역 | 공공데이터포털","url":"https://www.data.go.kr/data/15058207/openapi.do?recommendDataYn=Y"}]

Based on the search results, I found relevant information about the topics you searched for:

PublicDataReader 라이브러리

PublicDataReader는 공공 데이터를 자동으로 조회할 수 있는 파이썬 라이브러리입니다. 이 라이브러리로 공공데이터포털과 국가통계포털(KOSIS)과 같이 오픈 API 서비스로 제공하는 공공 데이터를 쉽게 조회할 수 있습니다. PublicDataReader를 이용하면 일반적인 공공 데이터 조회 과정에서 필요한 API 명세 찾기, 요청 작성, 반환된 데이터 정리 과정을 자동으로 처리할 수 있습니다.

식약처 의약품 데이터

식품의약품안전처에서 관리하는 의약품 품목, 주성분, 제조원, 포장단위, 저장방법, 성상등의 품목정보와 허가일자, 허가번호, 허가정보, 희귀의약품 여부 등의 허가받은 의약제품정보를 조회할 수 있는 OpenAPI가 공공데이터포털(data.go.kr)에서 제공되고 있습니다.

GitHub에서도 공공데이터(www.data.go.kr) - 식품의약품안전처 의약품 관련 정보-DUR품목정보(DURPrdlstInfoService)-병용금기정보조회(getUsjntTabooInfoList) 데이터 수집을 위한 저장소들이 있습니다.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[FETCH #44] https://raw.githubusercontent.com/jshchnz/claude-code-scheduler/main/examples/daily-review.json

ASK: Return the full JSON content verbatim.

{
  "version": 1,
  "tasks": [
    {
      "id": "daily-code-review",
      "name": "Daily Code Review",
      "description": "Review commits from the previous day for code quality and potential issues",
      "enabled": true,
      "trigger": {
        "type": "cron",
        "expression": "0 9 * * 1-5",
        "timezone": "local"
      },
      "execution": {
        "command": "Review all commits from yesterday. Check for: 1) Code quality issues, 2) Security vulnerabilities, 3) Performance concerns, 4) Missing tests. Summarize findings and suggest improvements.",
        "workingDirectory": ".",
        "timeout": 300
      },
      "tags": ["code-quality", "daily"],
      "createdAt": "2025-01-01T00:00:00.000Z",
      "updatedAt": "2025-01-01T00:00:00.000Z"
    }
  ],
  "settings": {
    "defaultTimezone": "local",
    "logRetentionDays": 30,
    "maxExecutionHistory": 100
  }
}

[FETCH #45] https://github.com/786raees/task-scheduler-python

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README usage (TaskScheduler class methods: create daily task, run, delete; win32com.client usage). Quote numbers precisely.

Task Scheduler Python Repository Summary

Repository Description: "The Task Scheduler Python project provides a convenient way to interact with the Windows Task Scheduler using Python and the win32com.client library."

Metrics:

  • Stars: 2
  • Forks: 0
  • Primary Language: Python

Last Commit: The repository shows "2 Commits" on the main branch, though a specific date isn't displayed in the provided content.

File Structure:

  • .vscode/ (configuration folder)
  • app/ (application folder)
  • .gitignore
  • LICENSE (MIT)
  • README.md
  • main.py
  • requirements.txt

Core Capabilities (TaskScheduler Class Methods):

The project enables developers to:

  • Create tasks using create_task() with executable path, arguments, and ISO 8601 trigger times
  • Retrieve tasks via get_all_tasks() for task enumeration
  • Toggle states using toggle_task() to enable/disable tasks
  • Execute tasks through run_task() method
  • Remove tasks via delete_task() functionality

Dependencies: Installation requires: pip install pywin32

License: MIT

[FETCH #46] https://github.com/jjscan/data.go.kr-1

ASK: Extract: repo description, star count, fork count, primary language, last commit date, file structure, README key contents (which data.go.kr MFDS API (DURPrdlstInfoService), how serviceKey is passed, pagination, output). Quote numbers precisely.

GitHub Repository Analysis: data.go.kr-1

Repository Metadata

Description: Collects drug interaction data from South Korea's public data portal (data.go.kr), specifically from the MFDS (Food and Drug Safety Ministry) DUR product information service's concurrent contraindication query API.

Metrics:

  • Stars: 0
  • Forks: 0
  • Primary Language: R

Last Commit: Not explicitly stated in provided content

File Structure

The repository contains two primary files:

  • DURPrdlstInfoService.R (main R code file)
  • README.md (documentation)

API Implementation Details

Target API: getUsjntTabooInfoList (병용금기정보조회 - Concurrent Contraindication Query)

Service: DURPrdlstInfoService (DUR Product List Information Service)

Data Source: "공공데이터(www.data.go.kr) - 식품의약품안전처 의약품 관련 정보"

How serviceKey is Passed

The documentation does not explicitly detail the serviceKey parameter mechanism. However, it references making GET requests to the public data API using the httr package.

Pagination

The API response structure indicates pagination parameters:

  • numOfRows: 100 (records per page)
  • pageNo: 895 (example page number shown)
  • totalCount: 351010 (example total records)

Output Format

Response data is in XML format, parsed using xmlSApply and xpathSApply functions. Data is ultimately exported to CSV files.

Key Implementation Challenges Documented

  1. Connection failures Resolved via polling mechanism with retry logic
  2. Null data handling Detected via totalCount == 0 validation
  3. R session crashes Unresolved intermittent issues
  4. Performance bottleneck Single-threaded processing required 17 hours for ~350,000 records; addressed through MPI parallel computing implementation using Rmpi

[FETCH #47] https://github.com/simonw/git-scraper-template

ASK: Extract: repo description, star count, fork count, last commit date, how the GitHub Actions workflow scrapes and commits (curl + git commit on change), files. Quote numbers precisely.

git-scraper-template Repository Summary

Description: "Template repository for setting up a new Git scraper using GitHub Actions."

Star Count: 132 stars

Fork Count: 10 forks

Last Commit Date: Not explicitly specified in the provided content

Workflow Mechanism: The GitHub Actions setup uses curl via a ./download.sh script to fetch content from a specified URL, then commits the result to the repository. The workflow runs on a configurable schedule (default: once every 24 hours).

Key Files:

  • .github/workflows/ (directory containing workflow configuration)
  • README.md
  • download.sh
  • scrape.sh (generated when creating a repository from this template)
  • .github/workflows/scrape.yml (workflow configuration file)

Optional Enhancement: The template supports Python-based scrapers via an uncommented block in scrape.yml paired with a requirements.txt dependency file.

[FETCH #48] https://www.fda.gov/drugs/drug-master-files-dmfs/list-drug-master-files-dmfs

ASK: Extract: the download link(s) for the DMF list file (xls/xlsx/zip), update cadence, columns described (DMF#, STATUS, TYPE, SUBMIT DATE, HOLDER, SUBJECT), and the "as of" date.

The server returned HTTP 404 Not Found.

The response body was not retrieved. If this URL requires authentication, use an authenticated tool (e.g. gh for GitHub, or an MCP-provided fetch tool) instead of WebFetch.

[FETCH #49] https://code.claude.com/docs/en/desktop-scheduled-tasks

ASK: Extract: what Claude Code Desktop scheduled tasks are, how they run (does the app need to be open?), Windows support, how they differ from routines and from claude -p with OS scheduler, and any limitations.

Documentation Index

Fetch the complete documentation index at: https://code.claude.com/docs/llms.txt Use this file to discover all available pages before exploring further.

Schedule recurring tasks in Claude Code Desktop

Set up scheduled tasks in Claude Code Desktop to run Claude automatically on a recurring basis for daily code reviews, dependency audits, or morning briefings.

Scheduled tasks start a new session automatically at a time and frequency you choose. Use them for recurring work like daily code reviews, dependency update checks, or morning briefings that pull from your calendar and inbox.

The Desktop app's Routines page lets you create both local scheduled tasks and remote routines. A local task runs on your machine with direct access to your files and tools, but only fires while the app is open and your computer is awake. A remote routine runs in the cloud even when your computer is off, and can also fire on API calls or GitHub events. This page covers local scheduled tasks; for remote routines and their trigger options, see Routines.

Compare scheduling options

Claude Code offers three ways to schedule recurring or one-off work:

Cloud Desktop /loop
Runs on Cloud, Anthropic-managed by default Your machine Your machine
Requires machine on No Yes Yes
Requires open session No No Yes
Persistent across restarts Yes Yes Restored on --resume if unexpired
Access to local files No (fresh clone) Yes Yes
MCP servers Connectors configured per task Config files and connectors Inherits from session
Permission prompts No (runs autonomously) Configurable per task Inherits from session
Customizable schedule Via /schedule in the CLI Yes Yes
Minimum interval 1 hour 1 minute 1 minute
Use **cloud tasks** for work that should run reliably without your machine. Use **Desktop tasks** when you need access to local files and tools. Use **`/loop`** for quick polling during a session. By default, scheduled tasks run against whatever state your working directory is in, including uncommitted changes. Enable the worktree toggle when creating the task to give each run its own isolated Git worktree, the same way [parallel sessions](/docs/en/desktop#work-in-parallel-with-sessions) work.

Create a scheduled task

On Claude Desktop before 1.1.5368, local scheduled tasks aren't available. In the Code tab, click Routines in the sidebar or in the sidebar's More menu, then click New routine and choose Local. Configure these fields:

Field Description
Name Identifier for the task. Converted to lowercase kebab-case and used as the folder name on disk. Must be unique across your tasks.
Description Short summary shown in the task list.
Instructions What Claude should do when the task runs. Write this the same way you'd write any message in the prompt box. The instructions input includes pickers for the permission mode and model, and below it you select the working folder and whether to run in an isolated worktree.
Schedule How often the task runs. See schedule options below.

A folder is required before you can save the task. If you haven't trusted that folder yet, Desktop prompts you to trust it before saving.

You can also create a task by describing what you want in any session. For example, "set up a daily code review that runs every morning at 9am" creates a recurring task, and "remind me at 3pm tomorrow to check the deploy" creates a one-time task that disables itself after it fires.

Schedule options

Pick a preset from the Schedule control:

  • Manual: no schedule, only runs when you click Run now. Useful for saving a prompt you trigger on demand
  • Hourly: runs every hour
  • Daily: shows a time picker, defaults to 9:00 AM local time
  • Weekdays: same as Daily but skips Saturday and Sunday
  • Weekly: shows a time picker and a day picker

For intervals the picker doesn't offer, such as every 15 minutes, the first of each month, or a single run at a specific future time, ask Claude in any Desktop session to set the schedule. Use plain language; for example, "schedule a task to run all the tests every 6 hours."

How scheduled tasks run

Scheduled tasks run on your machine. Desktop checks the schedule every minute while the app is open and starts a fresh session when a task is due, independent of any manual sessions you have open. Each task gets a small delay of a few minutes after the scheduled time to stagger API traffic. The delay is deterministic: the same task always starts at the same offset.

When a task fires, you get a desktop notification and a new session appears under a Scheduled section in the sidebar. Open it to see what Claude did, review changes, or respond to permission prompts. Claude can edit files, run commands, create commits, and open pull requests, the same as in a session you start yourself, but can't send or receive messages between your desktop sessions through the desktop app's session surface.

Tasks only run while the desktop app is running and your computer is awake. If your computer sleeps through a scheduled time, the run is skipped. To prevent idle-sleep, enable Keep computer awake in Settings under Desktop app → General. Closing the laptop lid still puts it to sleep. For tasks that need to run even when your computer is off, or that should trigger on an API call or GitHub event, create a remote routine instead.

Missed runs

When the app starts or your computer wakes, Desktop checks whether each task missed any runs in the last seven days. If it did, Desktop starts exactly one catch-up run for the most recently missed time and discards anything older. A daily task that missed six days runs once on wake. Desktop shows a notification when a catch-up run starts.

Keep this in mind when writing prompts. A task scheduled for 9am might run at 11pm if your computer was asleep all day. If timing matters, add guardrails to the prompt itself, for example: "Only review today's commits. If it's after 5pm, skip the review and just post a summary of what was missed."

Permissions for scheduled tasks

Each task has its own permission mode, which you set when creating or editing the task. Allow rules from ~/.claude/settings.json also apply to scheduled task sessions. If a task runs in Manual mode and needs to run a tool it doesn't have permission for, the run stalls until you approve it. The session stays open in the sidebar so you can answer later.

To avoid stalls, click Run now after creating a task, watch for permission prompts, and select "always allow" for each one. Future runs of that task auto-approve the same tools without prompting. You can review and revoke these approvals from the task's detail page.

MCP tools marked requiresUserInteraction prompt on every call and don't offer an always-allow option. Runs that call these tools stall each time.

Manage scheduled tasks

In the Code tab, click a task in the Routines list to open its detail page. From here you can:

  • Run now: start the task immediately without waiting for the next scheduled time
  • Status: toggle between Active and Paused to pause or resume scheduled runs without deleting the task
  • Edit: change the instructions, schedule, folder, or other settings
  • Review history: see every past run, including skipped runs. Hover a skipped entry to see why: your computer was asleep, the previous run was still in progress, or other scheduled tasks were already running. Click Show more to load older entries.
  • Review allowed permissions: see and revoke saved tool approvals for this task from the Always allowed panel
  • Delete: remove the task and archive all sessions it created. An Also delete files on disk checkbox appears in the confirmation dialog; check it to also remove the task's SKILL.md file and associated data from ~/.claude/scheduled-tasks/.

You can also list, create, edit, and pause tasks by asking Claude in any Desktop session. For example, "pause my dependency-audit task" or "show me my scheduled tasks." To delete a task, use the Delete button on its detail page.

A scheduled task can also modify its own schedule or prompt from within a running session using the update_scheduled_task MCP tool. This lets a task reschedule itself based on what it finds, for example, rescheduling a code review to run earlier when it detects a release branch has been created.

To edit a task's prompt on disk, open ~/.claude/scheduled-tasks/<task-name>/SKILL.md (or under CLAUDE_CONFIG_DIR if set). The file uses YAML frontmatter for name and description, with the prompt as the body. Changes take effect on the next run. Schedule, folder, model, and enabled state are not in this file: change them through the Edit form or ask Claude.

[FETCH #50] https://urlwatch.readthedocs.io/en/latest/filters.html

ASK: Extract the list of built-in filters (css, xpath, html2text, grep, strip, sort, re.sub, shellpipe, jq, etc.) with one example jobs.yaml snippet using css + html2text, and the diff-related options (diff_tool, diff_filter).

urlwatch Filters Summary

Built-in Filters

The documentation lists 28 built-in filters available in urlwatch:

beautify, css, csv2text, element-by-class, element-by-id, element-by-style, element-by-tag, format-json, grep, grepi, hexdump, html2text, pdf2text, pretty-xml, ical2text, ocr, re.sub, re.findall, reverse, sha1sum, shellpipe, sort, remove-duplicate-lines, strip, striplines, xpath, and jq.

Example: CSS + HTML2Text Filter

url: https://example.net/css.html
filter:
  - css: ul#groceries > li.unchecked
  - html2text

This configuration extracts specific HTML elements using CSS selectors, then converts the result to plaintext.

The documentation mentions two filter stages but does not document diff_tool or diff_filter options in this Filters page. It only references diff_filter briefly, stating it's "applied to the diff result before reporting the changes," and notes that --test-diff-filter can test these settings using cached historic data.

The page focuses on filters themselves rather than diff configuration options. For complete diff-related settings, the Configuration page would likely provide additional details.

[SEARCH #31] github 식약처 공고 크롤링 텔레그램 알림 스케줄러 파이썬 회수 판매중지

Web search results for query: "github 식약처 공고 크롤링 텔레그램 알림 스케줄러 파이썬 회수 판매중지"

Web search was not performed: this session has used its web search budget (200 of 200 WebSearch calls). Continue with the information already gathered instead of issuing more searches. If more searches are genuinely needed, ask the user to raise CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

Web search results for query: "github xlsxwriter multi-sheet report internal hyperlink summary sheet python generator repository"

Web search was not performed: this session has used its web search budget (200 of 200 WebSearch calls). Continue with the information already gathered instead of issuing more searches. If more searches are genuinely needed, ask the user to raise CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[FETCH #51] https://github.com/kirillkovalenko/nssm

ASK: Extract: repo description, star count, fork count, primary language, last commit date, whether it is the official NSSM source mirror, and latest version noted. Quote numbers precisely.

NSSM Repository Analysis

Repository Description: "Non-Sucking Service Manager"

Star Count: 1.2k stars

Fork Count: 169 forks

Primary Language: C++ (based on file extensions: .cpp, .h)

Last Commit Date: Not explicitly stated in the provided content

Official Source Mirror Status: This appears to be a community fork rather than the official source. The repository is maintained by user "kirillkovalenko" and the README references "http://nssm.cc/" as the primary documentation location, suggesting that is the official project site.

Latest Version Noted: Version 2.24, dated 2014-08-31 (from README header)

Key Detail: The README states NSSM "can start any application as an NT service and will restart the service if it fails for any reason," indicating this is a Windows service management utility written in C++.

[SEARCH #33] github Playwright python Windows Task Scheduler headless daily scrape report "pythonw" OR "schtasks" repository

Web search results for query: "github Playwright python Windows Task Scheduler headless daily scrape report "pythonw" OR "schtasks" repository"

Web search was not performed: this session has used its web search budget (200 of 200 WebSearch calls). Continue with the information already gathered instead of issuing more searches. If more searches are genuinely needed, ask the user to raise CLAUDE_CODE_MAX_WEB_SEARCHES_PER_SESSION.

REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.

[FETCH #52] https://github.com/mhammond/pywin32/issues/1563

ASK: Summarize the issue: what is requested regarding service recovery actions when a pywin32 service crashes, and any workaround mentioned (e.g., sc failure, ChangeServiceConfig2, exiting with non-zero / SvcStop not called).

Issue Summary

What's Requested: The user wants their win32serviceutil.ServiceFramework service to terminate in a way that triggers Windows recovery actions (like automatic restart). Normally, these actions activate when a service ends without reporting SERVICE_STOPPED status.

The Problem: When SvcRun() raises an exception or calls sys.exit(), pywin32's cleanup code automatically sets the service status to SERVICE_STOPPED, which prevents recovery actions from triggering.

Workaround Mentioned: The reporter currently uses os.kill(os.getpid(), signal.SIGABRT) to bypass the cleanup, but acknowledges this is an extreme measure they'd prefer to avoid.

Request: A mechanism allowing SvcRun() to flag that the automatic cleanup (specifically the SetServiceStatus() call) should be skipped, enabling recovery actions to execute as intended.

[FETCH #53] https://www.fda.gov/drugs/drug-master-files-dmfs

ASK: Find any link to the DMF list spreadsheet (xls/xlsx/zip download URL), the update cadence (quarterly), and the columns described. Quote URLs exactly.

The server returned HTTP 404 Not Found.

The response body was not retrieved. If this URL requires authentication, use an authenticated tool (e.g. gh for GitHub, or an MCP-provided fetch tool) instead of WebFetch.

[FETCH #54] https://www.data.go.kr/data/15095677/openapi.do

ASK: Extract: API name, endpoint base URL and operation names, request parameters, key response fields (ITEM_SEQ, ITEM_NAME, ENTP_NAME, ITEM_PERMIT_DATE, CANCEL_DATE, CANCEL_NAME, CHANGE_DATE, etc.), update cadence.

API Documentation Extract

API Information

API Name: 식품의약품안전처_의약품 제품 허가정보 (MFDS Medicine Product License Information)

Provider: 식품의약품안전처 (Ministry of Food and Drug Safety)

API Type: REST

Data Format: JSON + XML

Base Endpoint

The document does not explicitly specify the endpoint base URL.

Operations

The document references general query capabilities but does not list specific operation names or methods.

Request Parameters

Specific request parameters are not detailed in the provided content.

Key Response Fields

The API returns "의약품 품목, 주성분, 제조원, 포장단위, 저장방법, 성상등의 품목정보와 허가일자, 허가번호, 허가정보, 희귀의약품 여부 등의 허가받은 의약제품정보" (product items, active ingredients, manufacturers, packaging units, storage methods, appearance, license date, license number, rarity status).

The document mentions these potential fields: ITEM_SEQ, ITEM_NAME, ENTP_NAME, ITEM_PERMIT_DATE, CANCEL_DATE, CANCEL_NAME, CHANGE_DATE.

Update Cadence

Last modified: 2025-10-31. Specific update frequency is not specified.

Access Terms

  • Cost: Free
  • Usage Tier: Development (10,000 requests); Operations (expandable with use case registration)
  • Approval: Auto-approval for development; approval-based for operations

[FETCH #55] https://github.com/dgtlmoon/changedetection.io/wiki/Notification-configuration-notes

ASK: Extract the notification template tokens available (e.g. {{diff}}, {{diff_added}}, {{diff_removed}}, {{watch_url}}, {{current_snapshot}}, {{triggered_text}}) and an example notification body. Quote tokens exactly.

Notification Template Tokens

Based on the documentation, the following tokens are available for notifications:

  • {{base_url}}
  • {{current_snapshot}}
  • {{diff}}
  • {{diff_full}}
  • {{diff_added}}
  • {{diff_removed}}
  • {{watch_url}}
  • {{triggered_text}}

Example Notification Body

The guide provides a JSON notification example:

{
  'myKey': 1234,
  'url': '{{watch_url|tojson}}'
}

The documentation notes that {{current_snapshot}}, {{diff}}, and {{diff_full}} tokens can be lengthy and may exceed message length limits for certain services like Discord (2,000 characters). It recommends ensuring "your notification body contains at least something" to avoid delivery issues.

[FETCH #56] https://github.com/caronc/apprise/wiki/Notify_windows

ASK: Extract: the windows:// URL syntax, required dependencies (pywin32), parameters (duration, image), and limitations. Quote exactly.

Windows Notification with Apprise

URL Syntax

The basic syntax is straightforward: windows://

Required Dependencies

"pip install pywin32" is necessary to enable Windows notification functionality on your system.

Parameters

The documentation specifies one configurable parameter:

duration "Optionally set the duration of the popup message in seconds. By default this value is set to 12"

Example implementation: windows://?duration=5

Limitations

The key constraint is that "this notification can not be sent from one PC to another." Windows notifications only function when sending alerts to the same system you're currently using.

Message Specifications:

  • Icon support is available
  • Text message format
  • Maximum of 250 characters per message

[FETCH #57] https://github.com/jshchnz/claude-code-scheduler/tree/main/src

ASK: List the files/directories in src. Identify any file about Windows Task Scheduler (schtasks) or platform-specific schedulers and wrapper script generation.

Files/Directories in src

Based on the GitHub page listing, the src directory contains:

Subdirectories:

  • __tests__ (tests)
  • cron (cron scheduling)
  • history (history tracking)
  • logs (logging)
  • schedulers (scheduler implementations)
  • utils (utilities)
  • vcs (version control system)

Files:

  • config.ts
  • index.ts
  • types.ts

Windows Task Scheduler/Platform-Specific Findings

The provided content does not display file contents—only directory structure. Therefore, I cannot identify specific files related to Windows Task Scheduler (schtasks) or platform-specific scheduler wrappers from this listing alone.

However, the presence of a schedulers subdirectory suggests scheduler implementations may exist there. To locate Windows Task Scheduler or wrapper script generation code, you would need to examine the contents of the schedulers folder directly.

[FETCH #58] https://xlsxwriter.readthedocs.io/worksheet.html

ASK: Extract the exact signature and example of write_url() for internal links ('internal:Sheet2!A1'), add_table(), autofilter(), freeze_panes(), set_column(), and conditional_format() basics. Quote code.

XlsxWriter Method Signatures and Examples

Signature:

write_url(row, col, url[, cell_format[, string[, tip]]])

Internal Link Example:

# Link to a cell on the current worksheet.
worksheet.write_url('A1',  'internal:Sheet2!A1')

# Link to a cell on another worksheet.
worksheet.write_url('A2',  'internal:Sheet2!A1:B2')

# Worksheet names with spaces should be single quoted
worksheet.write_url('A3',  "internal:'Sales Data'!A1")

add_table()

Signature:

add_table(first_row, first_col, last_row, last_col, options)

Example:

worksheet.add_table('B3:F7', { ... })
# Or with row-column notation:
worksheet.add_table(2, 1, 6, 5, { ... })

set_column()

Signature:

set_column(first_col, last_col, width, cell_format, options)

Examples:

worksheet.set_column(0, 0, 20)   # Column A width set to 20
worksheet.set_column(1, 3, 30)   # Columns B-D width set to 30
worksheet.set_column('E:E', 20)  # Column E width set to 20
worksheet.set_column('F:H', 30)  # Columns F-H width set to 30

conditional_format()

Signature:

conditional_format(first_row, first_col, last_row, last_col, options)

Example:

worksheet.conditional_format('B3:K12', {'type':     'cell',
                                        'criteria': '>=',
                                        'value':    50,
                                        'format':   format1})

autofilter()

Not found in provided documentation.

freeze_panes()

Not found in provided documentation.

[FETCH #59] https://github.com/winsw/winsw/releases

ASK: Extract the latest release tag name and date, and the latest stable (2.x) release tag and date. Quote exactly.

WinSW Release Information

Latest Release:

  • Tag: "v3.0.0-alpha.11"
  • Date: "29 Jan 02:20"

Latest Stable (2.x) Release:

  • Tag: "v2.12.0"
  • Date: "28 Jan 16:22"

[FETCH #60] https://github.com/thp/urlwatch/releases

ASK: Extract the latest release tag name and date. Quote exactly.

Based on the webpage content provided, there is no latest release information available. The page displays:

"There aren't any releases here"

The repository currently has no published releases, so I cannot extract a release tag name or date.

[FETCH #61] https://github.com/dgtlmoon/changedetection.io/releases

ASK: Extract the latest release tag name and date. Quote exactly.

Latest Release Information

Tag Name: "0.55.8"

Release Date: "13 Jul 09:26"

This represents the most recent release of the changedetection.io project, featuring updates including Czech translation refinements, title extraction fixes, and various security and LLM-related improvements.

[FETCH #62] https://raw.githubusercontent.com/Q00/data.go.kr-crawling/master/url.py

ASK: Return the file content verbatim (URLs and API endpoints referenced).

Summary of Web Page Content

This Python script interacts with South Korea's public data portal API to extract parameter information.

Key Components

URL Referenced: https://www.data.go.kr/pubn/lab/gui/IrosDevGuide/selectReqResPrmList.do

Main Functionality: The script uses requests.post() to send JSON-formatted parameter data to the endpoint, then parses the response to extract field metadata.

API Parameters Used:

  • publicDataDetailPk: "uddi:9a60503c-b31c-4879-9028-a4250f0f6998"
  • paramtrSe: "2"
  • oprtinSeqNo: 15920

Processing Steps:

  1. Sends POST request with specified headers and JSON payload
  2. Parses JSON response to access RESULT_RE_LIST
  3. Iterates through result items (skipping index 8)
  4. Extracts parameter names and Korean translations
  5. Writes data structure to "column.py" file

Dependencies: The code imports gevent, base64, requests, BeautifulSoup, and json libraries to handle API communication and data transformation.

[FETCH #63] https://nedrug.mfds.go.kr/searchDmf

ASK: Describe this page: is it the 원료의약품 등록(DMF) 현황 search page? List the search form fields, result table columns, and pagination/parameters if visible.

Page Description

This is not the 원료의약품 등록(DMF) 현황 search page.

The page displayed is an error page from the Korea Food and Drug Administration's integrated pharmaceutical information system. It states: "The requested page cannot be found. The page you are looking for has been changed or is currently unavailable."

No search form fields, result table columns, or pagination parameters are visible on this page, as it is simply an error notification directing users to return to the main page via a link.

[FETCH #64] https://raw.githubusercontent.com/Q00/data.go.kr-crawling/master/go_data_crwaler.py

ASK: Summarize: which API endpoints it calls (data.go.kr MFDS?), how serviceKey/pagination is handled, how gevent is used, and output format. Quote key code lines.

API Usage Summary

API Endpoints: The script calls multiple endpoints from Korea's data.go.kr MFDS (Ministry of Food and Drug Safety) drug product information service:

  • getDurPrdlstInfoList (main product list)
  • getSeobangjeongPartitnAtentInfoList, getEfcyDplctInfoList, getOdsnAtentInfoList, getMdctnPdAtentInfoList, getCpctyAtentInfoList, getPwnmTabooInfoList, getSpcifyAgrdeTabooInfoList, getUsjntTabooInfoList (various attribute/warning endpoints)

ServiceKey & Pagination Handling: Authentication uses config.go_data_api_key. Pagination is managed manually: "params.update({'typeName' : column.typeName[addUrl]})" and "params_str2+= '&pageNo='+str(i+1)" construct paginated requests, with page count calculated as "page = int(totalCount/100) + 1".

Gevent: No gevent usage detected. The script uses standard synchronous requests.get() calls.

Output Format: Results are written to Excel files: "wb.save(column.typeName[addUrl]+'.xlsx')". Each endpoint generates a separate workbook with headers derived from XML response tags and data populated row-by-row.

[FETCH #65] https://github.com/huginn/huginn/wiki/Agent-Types

ASK: List agent types relevant to web change monitoring: WebsiteAgent, ChangeDetectorAgent, DeDuplicationAgent, DigestAgent, EmailDigestAgent, SchedulerAgent, ShellCommandAgent. Quote descriptions.

I cannot provide the requested information based on the content provided. The webpage shows a GitHub wiki page for Huginn's "Agent Types," but the actual content describing specific agent types and their functions failed to load properly (multiple "Uh oh! There was an error while loading" messages appear throughout).

While the page displays a JSON configuration example involving various workflow steps, it does not contain descriptions of the specific agent types you mentioned (WebsiteAgent, ChangeDetectorAgent, DeDuplicationAgent, DigestAgent, EmailDigestAgent, SchedulerAgent, ShellCommandAgent).

To obtain accurate agent type descriptions with proper quotes, you would need to access a working version of the Huginn documentation or wiki pages that successfully load the agent type reference material.

[FETCH #66] https://github.com/jshchnz/claude-code-scheduler/tree/main/src/schedulers

ASK: List the files in this directory (e.g., windows.ts, cron.ts, launchd.ts, base.ts). Quote names exactly.

Based on the webpage content, the files in the /src/schedulers directory are:

  1. "base.ts"
  2. "darwin.ts"
  3. "index.ts"
  4. "linux.ts"
  5. "windows.ts"

[FETCH #67] https://raw.githubusercontent.com/GitHub30/win11toast/main/README.md

ASK: Extract: dependencies (winsdk? winrt-*?), Python version requirement, examples for toast_async, notify(), buttons with on_click callback, app_id, and any note about running under Task Scheduler/service or without an interactive session.

win11toast Key Information

Dependencies

The documentation doesn't explicitly list package dependencies. However, the library is described as "based on WinRT," suggesting it relies on Windows Runtime capabilities. No specific winsdk or winrt-* package names are mentioned in the provided content.

Python Version

The page displays a PyPI badge for Python version support but doesn't specify the minimum required version in the text provided.

Examples for async/non-blocking usage

Non-blocking with notify():

from win11toast import notify
notify('Hello Python', 'Click to open url', on_click='https://www.python.org')

Async with toast_async():

from win11toast import toast_async
async def main():
    await toast_async('Hello Python', 'Click to open url', 
                     on_click='https://www.python.org')

Button with on_click callback

Buttons support protocol activation. Example: {'activationType': 'protocol', 'arguments': 'https://google.com', 'content': 'Open Google'} returns {'arguments': 'https://google.com', 'user_input': {}} when clicked.

App ID

No app_id parameter is documented in the provided content.

Task Scheduler/Service limitations

The documentation notes: "Since the current directory when executing the script is C:\Windows\system32, use os.chdir() accordingly." This implies constraints when running outside an interactive session, though explicit Task Scheduler/service guidance isn't provided.

[FETCH #68] https://github.com/huginn/huginn/blob/master/app/models/agents/website_agent.rb

ASK: Extract the agent description block: what WebsiteAgent does (scrape via css/xpath/json), 'mode' options (all, on_change, merge), 'schedule' default, 'extract' config example. Quote key text.

WebsiteAgent Overview

Core Functionality

The WebsiteAgent performs web scraping operations. As stated in the documentation: "The Website Agent scrapes a website, XML document, or JSON feed and creates Events based on the results."

Mode Options

Three operational modes are supported:

  • all: Creates events for every extraction result
  • on_change: Only generates events when data differs from previous results
  • merge: Retains existing payload data while updating with new values

Default Schedule

The agent runs on every_12h by default.

Extract Configuration Example

The documentation provides this sample configuration:

"extract": {
  "url": { "css": "#comic img", "value": "@src" },
  "title": { "css": "#comic img", "value": "@alt" },
  "hovertext": { "css": "#comic img", "value": "@title" }
}

This demonstrates CSS selector-based extraction where specific attributes (@src, @alt, @title) are pulled from HTML elements. The agent supports CSS selectors, XPath expressions, JSON paths, and regex patterns depending on document type.