Poetry Website Classification API Data Schema Explained

October 07, 2026
Poetry Website Classification API Data Schema Explained

You’re building a filter to control where content appears, and a big slice of your policy depends on whether a site is dedicated to poetry. By the end of this guide, you’ll integrate Klazify’s categorize endpoint, read the returned fields that indicate poetry-related content, and wire the results into actions like blocking, enrichment, or auditing in your pipeline.

Why Klazify is the best fit for poetry-specific site classification

Poetry content can be sparse, stylized, and embedded in broader arts pages, which makes it harder to classify using basic metadata alone. Klazify’s models analyze on-page content and signals from the entire site, so your logic can distinguish a poetry archive from a general book retailer, a literary magazine’s poetry section from news coverage, or a poet’s portfolio from a school curriculum page.

Illustration: Poetry Website Classification API Data Schema Explained
  • Accurate website categorization using AI: Klazify analyzes actual page content beyond titles or meta tags, improving precision on niche literary content such as poem collections, poet bios, and submission guidelines.
  • Global coverage: Poetry sites are multilingual; the classifier handles multiple languages so you can evaluate small-press and international poetry hosts consistently.
  • Real-time classification: Get current results rather than static lists, so freshly published poem pages and modern poetry journals are captured as they evolve.
  • Industry-level categories: Mapping to the IAB taxonomy helps your ad and safety tooling align with industry-standard categories for arts and literary content relevant to poetry.
  • Simple API integration: A single REST call provides categories, company fields, logo, technology hints, registration data, and more, which is useful when poetry sites lack extensive structured data.
  • Compliance and filtering: Use the category names and confidence scores to enforce allowlists and blocklists wherever poetry content should or should not appear.

Scenario: block, enrich, or audit poetry destinations with one call

Common operational goals around poetry content:

  • Brand safety and placement controls: Exclude poetry-focused destinations when your campaign targets different interests, or include them when poetry is in-scope.
  • Signup enrichment: When a user adds a domain, label the account with “Poetry-related” if the content categories match your mapping, then assign to the correct team or product tier.
  • Inventory auditing: Nightly scans of partner lists ensure poetry sites are correctly tagged for routing across your network.

All of the above start with a call to the categorize endpoint and then using the returned category paths and confidence values to match your policy.

The single endpoint you need

Endpoint: POST https://www.klazify.com/api/categorize

  • Authorization: Bearer token in the Authorization header.
  • Body: JSON with a url field.
  • Response: Category list with confidence, optional IAB mapping codes when available, logo URL, company object, domain registration data, and a list of similar domains.

Notes for planning:

  • Batched or parallel requests: You can parallelize within your client according to your throughput goals. Observe your plan’s limits; consult the Documentation for current guidance.
  • Caching: Cache per domain. Poetry sites change slowly; conservative TTLs reduce cost and latency. Refresh more frequently for publications with daily updates.
  • Error policy and billing: Failed or unreachable calls are not billed. Implement retries with backoff and move on if the host is consistently unreachable.
  • Plan info: Starter is $39.99/mo with a 7-day trial.

Make your first request

Use the official docs fixture below to verify your pipeline and auth. Replace YOUR_API_KEY with your token.

# Label: curl request
curl -X POST "https://www.klazify.com/api/categorize" \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://cbsnews.com"}'

Understand the response fields you’ll rely on

Here is the official example JSON. Use it to validate parsing and to build your field mapping. Do not modify field names when you code against the live API.


{
"domain": {
"categories": [
{
"confidence": 0.92,
"name": "/Computers & Electronics/Consumer Electronics",
"IAB-632-596": "Consumer Electronics/Technology & Computing/Consumer Electronics"
},
{
"confidence": 0.89,
"name": "/Internet & Telecom/Mobile & Wireless/Mobile Phones"
}
],
"social_media": null,
"logo_url": "https://klazify.s3.amazonaws.com/2110787991611585019600ed5fb1d1300.04730104.png"
},
"success": true,
"objects": {
"company": {
"url": "https://www.apple.com/",
"name": "Apple",
"city": "Cupertino",
"stateCode": "CA",
"countryCode": "US",
"employeesRange": "100K+",
"revenue": 274515000000,
"raised": null,
"tags": [
"E-commerce",
"Consumer Electronics",
"Mobile",
"B2C"
],
"tech": [
"omniture_adobe_analytics",
"atlassian_confluence",
"successfactors",
"apache_apex",
"talend",
"oracle_peoplesoft",
"salesforce",
"stripe",
"dell_boomi_atomsphere",
"gigya",
"sage_50cloud",
"quickbooks",
"webmethods",
"apache_tomcat",
"alteryx",
"tibco_rendezvous",
"atlassian_jira",
"..."
]
}
},
"domain_registration_data": {
"domain_age_date": "1987-02-19",
"domain_age_days_ago": "13026",
"domain_expiration_date": "2030-02-20",
"domain_expiration_days_left": "123"
},
"similar_domains": [
"bestbuy.com",
"icloud.com",
"microsoft.com",
"macrumors.com",
"google.com",
"samsung.com",
"twitter.com",
"hp.com",
"bhphotovideo.com",
"dell.com"
]
}

How to use the fields:

  • domain.categories: Each entry includes a category path in name and an optional IAB-… mapping. Use confidence to set thresholds (for example, require at least one category over a chosen value to tag as poetry-related). You will maintain your own list of category paths that indicate poetry content and compare against name values.
  • domain.logo_url: Fetch a brand logo for UI labels on poetry-related sites, or to show reviewers what was categorized.
  • objects.company: Enrichment source when a poetry site represents an organization (e.g., a literary magazine). Fields like name, city, countryCode, employeesRange, revenue, tags, tech can inform routing or scoring. Treat revenue as a large integer; do not assume currency in code without confirming in the Documentation.
  • domain_registration_data: Dates are strings in YYYY-MM-DD; interpret them as calendar dates (UTC). Useful for heuristics (e.g., deprioritize very new domains).
  • similar_domains: For discovery and QA: if a domain is poetry-related, neighboring domains often share similar editorial focus.

Code: classify and map to your “Poetry” tag

The snippet below calls the endpoint, extracts categories, and maps to a custom Poetry label using your own list of category paths that you consider poetry-related. Persist the decision along with the raw JSON for audits.

# Label: Python script
import json
import requests
from typing import List, Dict

API_URL = "https://www.klazify.com/api/categorize"
API_KEY = "YOUR_API_KEY"

# Maintain this list in your config store.
# It should contain the category path strings from 'domain.categories[].name'
# that you consider poetry-related.
POETRY_CATEGORY_PATHS = set([
# e.g., "/Arts & Entertainment/..." or other paths you curate internally.
# Populate from observed responses; keep it versioned.
])

def is_poetry(categories: List[Dict], threshold: float = 0.6) -> bool:
for cat in categories:
name = cat.get("name")
conf = float(cat.get("confidence", 0))
if name in POETRY_CATEGORY_PATHS and conf >= threshold:
return True
return False

def categorize(url: str) -> Dict:
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
payload = {"url": url}
r = requests.post(API_URL, headers=headers, json=payload, timeout=20)
r.raise_for_status()
return r.json()

def main():
url = "https://cbsnews.com" # Replace with the domain under review
data = categorize(url)
domain = data.get("domain", {})
categories = domain.get("categories", []) or []

poetry_flag = is_poetry(categories, threshold=0.7)

result = {
"url": url,
"poetry_related": poetry_flag,
"categories": [(c.get("name"), c.get("confidence")) for c in categories],
"logo_url": domain.get("logo_url"),
"company_name": (data.get("objects", {}) or {}).get("company", {}).get("name"),
"domain_age_date": (data.get("domain_registration_data", {}) or {}).get("domain_age_date"),
"similar_domains": data.get("similar_domains", []),
}

print(json.dumps(result, indent=2))

if __name__ == "__main__":
main()

Implementation notes:

  • Confidence thresholding: Start conservatively (e.g., 0.7) and adjust after sampling. Retain the raw category list for reviewers.
  • Null handling: social_media can be null; always check for None before iterating.
  • Type safety: revenue can exceed 32-bit range; store as 64-bit integer if you persist it.

Building your poetry taxonomy mapping

Because editorial sites vary, you should define a mapping set of category paths that map to your internal Poetry label. You obtain these paths from domain.categories[].name values observed in sample responses and refine over time.

  • Start with a seed list: Run the endpoint on a curated set of known poetry magazines, poem archives, and poet portfolios. Log all category paths and their average confidence.
  • Create a mapping set: Select the category paths that consistently appear on poetry sites in your dataset.
  • Set thresholds by path: For strong signals, use a higher confidence requirement; for broader arts paths, require multiple category matches.
  • Review drift: Re-sample quarterly to capture new category paths that emerge as the web changes.

Which fields drive your decision, and how to act

Field Role in poetry decision Common action Caching guidance
domain.categories[].name + confidence Primary signal: compare names against your Poetry mapping, apply thresholding Set poetry_related flag; route to block/allow; tag account Cache 7–30 days for stable sites; 1–7 days for frequently updated publishers
objects.company.tags Secondary hint: presence of relevant business tags can support a borderline decision Use as tie-breaker when category confidence is near threshold Same as categories; refresh with category refresh
domain_registration_data.domain_age_date Quality heuristic: very new domains may need cautious handling Hold for manual review or require higher confidence Cache long-term; dates change infrequently
similar_domains[] Discovery: find additional poetry-related properties to scan Queue neighbors for classification Refresh periodically; list can evolve with content
domain.logo_url UX: provide a recognizable brand image during review Render in dashboards and audit logs Cache indefinitely; update on next scheduled refresh

Operational guidance: throughput, reliability, and unknowns

  • Batching and concurrency: Group domains and send in parallel to meet your SLA. Coordinate workers to avoid duplicate lookups for the same domain within your cache TTL.
  • Retry strategy: For timeouts or temporary DNS failures, retry with exponential backoff. Because failed or unreachable calls aren’t billed, it’s safe to implement cautious retries without cost spikes.
  • Handling new or unknown domains: If the response lacks confident category matches, store poetry_related as false and mark for rescan later. You can also add a second policy requiring multiple positive signals before allowing a new domain.
  • Consistency: Normalize domains (lowercase, strip protocol, remove trailing slashes) before caching keys to prevent duplicate entries.
  • Auditing: Store the raw JSON response alongside your computed flags. This makes it easy to explain why a site was classified as poetry-related.

Security, platform endpoints, and MCP

  • Authentication: Use the Bearer token in the Authorization header. Never embed secrets in client-side code.
  • MCP note: The MCP host exists at https://mcp.klazify.com. A GET to /mcp returns HTTP 405. Use the categorize endpoint described above for classification. See the MCP page for details.
  • Docs and onboarding: Review request/response schemas and implementation notes in the Documentation. Create a free account to obtain your API key and start the 7-day trial.

Putting it into production: policy examples for poetry

  • Blocklist policy: If any category path from your Poetry mapping is present with confidence ≥ 0.7, block placements on that domain and log the matched path(s).
  • Enrichment policy: When a new customer signs up with a domain identified as poetry-related, add an internal “Poetry” attribute, set the right product tier, and assign a specialist CSM.
  • Audit policy: For networks that should include poetry destinations, run a nightly scan; alert when a previously poetry-related site no longer matches your mapping (e.g., site pivoted content).

End-to-end checklist

  • Normalize domains and deduplicate.
  • Call POST /api/categorize with Bearer auth and JSON body.
  • Parse domain.categories, compare against your Poetry mapping, and apply thresholding.
  • Store raw JSON and computed flags; cache results.
  • Execute actions (block, allow, tag, audit) and schedule refreshes.

FAQ

  • Which endpoint and method classify a domain for poetry-related content?
    Use POST https://www.klazify.com/api/categorize with a JSON body containing {"url": "..."} and a Bearer token.
  • Are failed or unreachable calls billed?
    No. Failed/unreachable calls are not billed, so you can safely retry with backoff.
  • How should I cache results for poetry sites?
    Cache per domain. For stable literary archives, a longer TTL reduces latency and cost. For frequently updated publications, shorten the TTL and schedule refreshes.
  • What if the response doesn’t clearly indicate poetry?
    Treat the domain as not poetry-related by default, queue for periodic rescans, and require stronger evidence (multiple category hits or higher confidence) before changing status.
  • How do I align Klazify categories with my internal “Poetry” label?
    Collect category paths from your seed list of poetry domains, build a mapping set from domain.categories[].name, and maintain it over time. Use confidence thresholds to reduce false positives.

Ready to ship your integration? Create your account to get an API key and start the 7-day trial: Register. Explore fields, examples, and additional guidance in the Documentation. For platform notes, see MCP.

Ready to use Klazify?

Start classifying websites, enriching company data, and exploring web intelligence.

Get Started Free