Poetry URL Classification API: Technical Implementation Notes
You need to automatically detect whether a URL is about poetry so you can enforce filtering in schools, enrich new writer signups, or audit where your ads are running. By the end of this guide, you’ll have a working, production-ready path to classify poetry-related URLs using Klazify’s POST /api/categorize endpoint, parse the response, and wire it into your pipeline with caching, error handling, and taxonomy mapping.
Why Klazify is a strong fit for poetry-focused URL classification
Poetry websites range from minimalist author pages to digital magazines and literary journals with mixed content like reviews, contests, and events. Klazify’s feature set is well-suited to separate poetry content from broader arts content at scale:
- Accurate website categorization using AI: The API analyzes full-page content, which is critical when the word “poetry” isn’t in the domain but the page body, titles, or navigation reveal it.
- Global coverage: Many poetry journals and blogs are multilingual; Klazify’s analysis works across languages so you can classify international literary sites.
- Real-time classification: Poetry magazines publish new issues frequently. Real-time analysis helps you avoid stale allow/deny decisions.
- Industry-level categories: Results are mapped to the IAB taxonomy, letting you align poetry-related decisions with standard ad tech and brand safety labels.
- Simple API integration: One REST endpoint to get categories, company data, social signals, logo, and similar domains—useful for both filtering and enrichment.
- Compliance and filtering: Strong signals for whitelisting poetry resources in education or excluding them from campaigns that shouldn’t appear on literary content.
The concrete scenario: detect and act on poetry URLs
Let’s anchor on three common workflows:
- Content filtering for education networks: Allow poetry resources but block unrelated categories; log ambiguous results for review.
- Ad placement and brand safety: Include poetry publications in a specific contextual segment while excluding unrelated arts content.
- Signup enrichment: When a new author adds a portfolio URL, classify it and attach IAB categories, logo_url, and company tags to your CRM profile.
Each workflow uses the same endpoint: send a URL, parse domain.categories, map the IAB taxonomy codes/names to your internal categories, then apply your allow/deny or enrichment logic.
Endpoint and authentication
Main endpoint:
- URL: https://www.klazify.com/api/categorize
- Method: POST
- Auth: Bearer token in the Authorization header
- Body: JSON with a single field “url”
Pricing notes relevant to rollout: Starter $39.99/mo with a 7-day trial; failed or unreachable calls are not billed. For full parameter details, see the Documentation.
Quickstart: one cURL to classify a URL
Request (copy/paste and replace YOUR_API_KEY):
curl -X POST "https://www.klazify.com/api/categorize" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://cbsnews.com"}'
Notes:
- Use a fully-qualified URL (including protocol) for best results.
- You can post any poetry-related URL (e.g., a journal’s issue page) to classify at the URL level, not just domain level.
- Unreachable URLs won’t be billed.
Response schema and the fields you’ll use
The example below shows the structure you’ll receive. Use it to understand which fields will power your poetry workflows. Do not rely on these specific values for your own URLs—treat them as a schema reference.
{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}
How to use these fields in a poetry pipeline:
- domain.categories: Your primary signal. Inspect name and any IAB mapping fields. Compare them to your poetry allowlist/denylist rules. Use confidence as a threshold for automated decisions vs. manual review.
- domain.logo_url: Helpful for UI moderation dashboards or CRM enrichment when a poet adds a site.
- objects.company: Use tags and name when you need organizational context (e.g., a literary magazine vs. a personal blog that lacks a formal company record).
- domain_registration_data: Optional quality signal for trust or additional scoring.
- similar_domains: Useful for discovery—queue these for background checks to find more poetry-related sites to include or exclude.
- success: Always validate this before reading other fields.
Python example: classify and route poetry URLs
import os
import requests
from typing import Dict, Any, List, Optional
API_URL = "https://www.klazify.com/api/categorize"
API_KEY = os.getenv("KLAZIFY_API_KEY", "YOUR_API_KEY")
def classify_url(url: str) -> Dict[str, Any]:
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json"
}
payload = {"url": url}
r = requests.post(API_URL, headers=headers, json=payload, timeout=20)
r.raise_for_status()
return r.json()
def is_poetry(categories: List[Dict[str, Any]], poetry_allowlist: List[str], min_conf: float = 0.5) -> bool:
# poetry_allowlist contains category names or IAB mappings you consider "poetry" or "literature"
for c in categories or []:
name = c.get("name") or ""
conf = float(c.get("confidence") or 0.0)
if conf < min_conf:
continue
# match either the readable name or any present IAB mapping value
if any(allow in name for allow in poetry_allowlist):
return True
for k, v in c.items():
if k.startswith("IAB-") and isinstance(v, str):
if any(allow in v for allow in poetry_allowlist):
return True
return False
def route_decision(resp: Dict[str, Any], poetry_allowlist: List[str]) -> str:
if not resp.get("success"):
return "retry_or_manual_review"
domain = resp.get("domain") or {}
categories = domain.get("categories") or []
if is_poetry(categories, poetry_allowlist):
return "allow_or_include_segment_poetry"
# if confidence too low or mixed signals, send to review
if any((c.get("confidence") or 0) < 0.5 for c in categories):
return "manual_review"
return "deny_or_exclude"
if __name__ == "__main__":
test_url = "https://example.com/poems/latest-issue"
resp = classify_url(test_url)
# Configure your own allowlist entries that represent poetry-related values
poetry_allowlist = [
# Example substrings you match against category names or IAB mappings
# e.g., "Literature", "Poetry" if present in your taxonomy
]
decision = route_decision(resp, poetry_allowlist)
# Optional enrichment for CRMs or moderation UIs
logo_url = (resp.get("domain") or {}).get("logo_url")
company = (resp.get("objects") or {}).get("company") or {}
company_name = company.get("name")
tags = company.get("tags") or []
similar = resp.get("similar_domains") or []
print("Decision:", decision)
print("Logo:", logo_url)
print("Company:", company_name, "Tags:", tags)
print("Similar domains:", similar)
Mapping the response to your poetry taxonomy
Many teams already have an internal taxonomy that includes poetry as a leaf node or as part of a broader “literature” branch. Use the following approach to map Klazify results to your taxonomy without hard-coding vendor-specific labels:
- Build a translation table keyed by category name and any available IAB mapping fields returned in domain.categories[].
- For ambiguous sites (e.g., publications mixing essays, fiction, and poetry), set a confidence threshold and capture multiple categories to determine a final label.
- Log unknown categories and update your translation table periodically.
| Field | What it represents | How to use for poetry |
|---|---|---|
| domain.categories[].name | Readable category path | Pattern-match against your poetry-related labels; treat as primary signal. |
| domain.categories[].confidence | Score for the category | Set decision thresholds; lower scores go to manual review. |
| domain.categories[].IAB-* | IAB taxonomy mapping | Map to internal taxonomy if you align with IAB for ad tech. |
| objects.company.tags | Business and topical tags | Secondary hints; enrich profiles of poetry publishers or authors. |
| domain.logo_url | Logo image URL | Display in moderation dashboards and CRM records. |
| similar_domains[] | Related domains | Queue discovery of more poetry sites for your allowlist. |
Operational guidance for production
Batching and throughput
- Batch domain processing by queueing URLs and making concurrent POST requests within your service’s limit. Consult the Documentation for the latest guidance on rate limits.
- If you have multi-URL ingestion (e.g., crawling a poetry magazine’s sitemap), enqueue individual pages for URL-level classification when you need fine-grained context.
Caching and revalidation
- Cache results per canonical domain and per URL. Poetry homepages may be stable, but issue pages can change as new poems are added.
- Choose TTLs that reflect how often the content updates. On a cache hit, only revalidate if the page has changed (e.g., using your own last-seen hash or fetch timestamp heuristic).
- Store the full response together with the timestamp and the decision you made to aid auditing and reproducibility.
Error handling and retries
- Check success before reading fields. If false or absent, use backoff and retry logic; unreachable calls are not billed.
- For network issues or timeouts, place the URL back on the queue with a capped retry count and mark as “manual_review” if exhausted.
Handling unknown or new poetry domains
- When categories are empty or confidence is low, flag the domain for human review. Store reviewer overrides in your mapping table.
- Leverage similar_domains to schedule background checks that can uncover additional poetry sites from a trusted source.
Security and compliance considerations
- Log only the minimum necessary response fields for compliance. If you store company fields, secure PII-adjacent data like location.
- Use HTTPS everywhere and keep API keys in a secure secret manager.
Decision patterns for common poetry use cases
Content filtering on school networks
- Allowlist: Maintain a list of poetry-oriented category names and IAB mappings that you approve.
- Confidence gates: Allow automatically if any poetry-matching category exceeds your threshold; otherwise, route for review.
- Mixed content: If a URL belongs to a general arts site, prefer URL-level classification for the poems section over domain-level rules.
Ad placement and brand safety
- Context inclusion: Build a segment from URLs/domains with poetry-matching categories. Attach IAB mappings for interoperability with your ad stack.
- Exclusions: Maintain a separate denylist for unrelated categories that sometimes co-occur on literary magazines (e.g., unrelated events or product reviews).
- Audit: Store classification snapshots with decision outcomes to verify where creatives ran.
Signup enrichment for author platforms
- When a user provides a portfolio link, classify it and attach category names, confidence, logo_url, and company tags to their profile.
- Use similar_domains to suggest additional verified links the author might own (e.g., affiliated literary journals) and improve discovery.
Quality assurance and monitoring
- Sampling: Periodically sample poetry decisions and compare to human judgments; tighten or relax your thresholds accordingly.
- Drift checks: Track distribution changes in category names; spikes may indicate taxonomy drift, crawling anomalies, or content shifts.
- Override feedback loop: Feed manual reviewer overrides back into your mapping table to reduce future ambiguity.
MCP and ecosystem notes
If you’re evaluating management control plane options or exploring ecosystem integrations, see the MCP page. The address https://mcp.klazify.com exists; a direct GET on /mcp responds with 405, which is expected behavior and not an API you need for classification. Your poetry integration relies only on POST https://www.klazify.com/api/categorize.
Putting it all together
Implement the following minimal flow:
- Collector: Accept URLs from users, crawlers, or feeds.
- Classifier: POST each URL to /api/categorize and validate success.
- Mapper: Compare domain.categories entries and any IAB mappings to a maintained poetry allowlist, using confidence thresholds.
- Decision: Allow/deny for filtering, include/exclude for ads, or enrich for CRM with logo_url and company fields.
- Cache: Save results and decisions with timestamps; schedule revalidation.
- Monitor: Track error rates, drift in categories, and manual override volume.
FAQ
- Can I classify specific poem pages, not just the site? Yes. Send the full URL in the url field. URL-level classification is useful when only a subsection of a site contains poetry.
- How do I avoid re-billing on the same domain? Cache results per domain and per URL with your own TTLs. Reuse cached labels when appropriate and revalidate on a schedule or on content change.
- What should I do if the response has low confidence or no categories? Route the URL to manual review, log the outcome, and add a mapping override for future automation. You can also queue similar_domains for discovery and triangulation.
- How do I align Klazify categories with my internal taxonomy? Maintain a translation table that maps domain.categories[].name and any IAB mapping fields to your taxonomy nodes. Update it as you encounter new labels.
- What happens if the URL is down? Unreachable calls are not billed. Implement retries with backoff and mark exhausted attempts for manual review.
Ready to classify poetry URLs at scale? Create your account to get an API key and start integrating today: Register. For more details on fields and response formats, see the Documentation.
Ready to use Klazify?
Start classifying websites, enriching company data, and exploring web intelligence.
Get Started Free