Saltar para o conteúdo

Notícias e artigos

Trace every story from discovery to the page you actually observed.

Discover public news results through the documented Google Search API news vertical, access eligible public article pages, or scope a maintained feed that keeps publication, article, revision, and provenance records distinct.

  1. 01
    Name the object

    Discovery result, publication, article, author, revision, or topic.

  2. 02
    Keep the evidence

    Query context, source URL, source fields, and observation time.

  3. 03
    Set the boundary

    Metadata, eligible body fields, change history, and delivery ownership.

Story record model

One news story can appear as several different records.

Model discovery, publication, content, people, revisions, and subjects separately so a result snippet never masquerades as a complete article.

01 · DiscoveryDocumentado

News result

A query-bound result shown in the Google News vertical.

  • Headline and destination
  • Source and snippet
  • Query and request context
02 · SourceDocumentado

Publication

The public publisher identity and source page associated with an observation.

  • Source name and URL
  • Section or channel
  • Language and locale
03 · ContentPrimeiro piloto .

Article

Source-linked headline, byline, body, metadata, and media references where the eligible page exposes them.

  • Canonical URL
  • Visible article fields
  • Author or byline
04 · ChangePrimeiro piloto .

Edition or revision

A successive observation of selected article fields, never an inferred editorial intent.

  • Content hash
  • First and last seen
  • Changed field set
05 · MeaningPrimeiro piloto .

Topic or entity

An enrichment layer produced under agreed models or rules and kept separate from source facts.

  • Named entity cue
  • Topic label
  • Rule or model version
Object boundaryNot assumed

A discovery result is not a complete article.

Discovery is not the same as full article extraction. A snippet, article page, author byline, publisher, revision, and topic each carry a different evidence boundary.

Coverage contract

Separate documented access from source-specific normalization.

Coverage is confirmed by discovery surface, article source, page family, locale, required fields, and responsible-use boundary.

Engine or familyIntroduçãoContextoEstado
Google News através da API de Pesquisa do GoogleQuery with tbm=nws and supported request contextVisible discovery resultsDocumentado
Páginas de artigos públicos elegíveisURL submitted to Scraper API or Browser APIHTML or rendered page for your extraction workflowDocumentado
Alimentos normalizados de artigosApproved source list and field contractSource-agnostic schema, dedupe, topics, publisher feedsPrimeiro piloto .
Conteúdo restritido e reedição geralPaywall, login, membership, or unrestricted body reuseOutside standard public-web scopeNão padrão

Record schema

Make source claims, derived fields, and time semantics inspectable.

Keep what the source exposed separate from what the collection observed and what a downstream enrichment inferred.

01 · Discovery

Result context

query
The submitted discovery query.
rank
Position in this result page and request context.
result_url
The destination exposed by the result.
02 · Article

Source-visible fields

headline
Headline exposed on the eligible page.
byline
Public author string, not automatically a resolved person.
body_state
Present, absent, excluded, failed, changed, or review.
03 · Time

Publication and observation

published_at
Time attributed by the source, when exposed.
observed_at
Time the collection produced this observation.
first_seen_at
First appearance within the contracted monitoring window.
04 · Provenance

Traceability and state

source_url
Public page associated with the record.
content_hash
Optional fingerprint for successive observations.
schema_version
Version of the agreed output shape.
Illustrative article record—not customer dataJSON
{
  "record_id": "story_0042",
  "object_type": "article_observation",
  "discovery": {
    "query": "energy outlook",
    "surface": "google_news",
    "rank": 3
  },
  "article": {
    "headline": "Grid investment accelerates",
    "publication": "Example Daily",
    "body_state": "observed"
  },
  "published_at": "YYYY-MM-DDT06:10:00Z",
  "observed_at": "YYYY-MM-DDT08:42:00Z",
  "source_url": "https://news.example/story/42",
  "schema_version": "news.v1"
}
Two clocks, two meanings.

published_at is source-attributed publication time; observed_at is collection time. Neither should overwrite the other.

Identity and lineage

Group related records without erasing each source.

Preserve the publisher, URL, discovery context, and article observation before adding canonical URLs, clusters, entities, or topic labels.

  1. 01

    Discovery key

    query + result URL

    Request context retained.
  2. 02

    Source cues

    canonical URL + publisher

    Used when exposed and verified.
  3. 03

    Content state

    headline + hash + time

    Compared under a named rule version.
  4. 04

    Relationship

    same article, revision, or candidate cluster

    Ambiguity remains explicit.

Freshness semantics

A recently published story and a recently observed page are different facts.

Set collection cadence by workflow, then retain the time attributed by the source alongside request, observation, delivery, and change times.

  1. Source timepublished_at

    The displayed publication time, exactly as interpreted under the contract.

  2. Collectedobserved_at

    The article or result was observed in the requested context.

  3. Delivereddelivered_at

    The record became available to the customer.

  4. Changedfirst_seen / last_seen

    The observation entered or left the monitored history.

Quality and missingness

Do not turn an unavailable article into an empty story.

Carry the state that explains why a field is missing so discovery, parsing, source availability, and policy exclusions remain distinguishable.

Observado

Required structure present

The record passed agreed discovery or article field checks.

Source absent

Field not exposed

The source page did not show the requested byline, time, or body field.

Revisão

Template or lineage is ambiguous

The page or relationship needs inspection under the contracted rules.

Unavailable

No supported observation

Excluded, restricted, failed, or unavailable remains distinct from an empty article.

Structural checks

Required keys, types, result counts, body-state reason, timestamp shape, and schema version.

Field checks

URL format, source attribution, publication-time parse state, language, and optional content hash.

Acceptance boundary

Your team approves representative records, source set, permissible use, and decision thresholds; we operate the agreed technical checks.

Operating model

Choose where access ends and the maintained record begins.

Compare the same story lifecycle across infrastructure, APIs, recurring feeds, and a managed data operation.

Maximum control

Build the full news workflow on proxy infrastructure.

Your team chooses discovery surfaces, collects eligible pages, extracts article fields, resolves lineage, monitors changes, and delivers records.
ResponsabilidadeProprietário
Source and field briefCustomer
Access infrastructureShared
Extraction and schemaCustomer
Quality and maintenanceCustomer
Storage and decisionsCustomer

Applications

Use one evidence model across discovery, monitoring, and research.

The record layer supplies source-linked observations; your team defines analytical models, materiality, thresholds, and decisions.

01

Media monitoring

Track named queries, sources, and public story fields over time.

Results · articles · revisions
02

Market and topic research

Build reviewable corpora for trend, narrative, and source analysis.

Articles · topics · provenance
03

AI retrieval and grounding

Prepare time-aware, source-linked inputs for approved retrieval workflows.

Text state · source URL · timestamps
04

Reputation and risk signals

Surface new public coverage for human review without presenting it as a conclusion.

Discovery · entities · review state
05

Publisher intelligence

Compare public publishing patterns, sections, and revision behavior.

Publication · article · first seen
06

Research archives

Retain agreed metadata and permissible content with traceable collection history.

Schema version · provenance · change history

Representative pilot

Prove discovery, extraction, lineage, and missing states together.

Use real query contexts and representative public article pages, including ordinary stories, revisions, absent fields, changed templates, and restricted states.

Scope a news data pilot
  1. 01

    Define

    Queries, sources, article fields, contexts, cadence, and permitted use.

  2. 02

    Sample

    Results and pages that represent source templates, missing fields, and changes.

  3. 03

    Inspect

    Source evidence, times, body states, lineage candidates, and exceptions.

  4. 04

    Accept

    Approve the schema, source set, operating model, and technical thresholds.

Evaluation FAQ

Clarify the article boundary before scaling collection.

Answers describe documented capabilities and the terms that a source-specific pilot must confirm.

What news data is documented today?

The Google Search API supports the Google News result vertical through the tbm=nws parameter. It is suited to discovering visible news results for a defined query and request context. Eligible public article pages can also be collected through the Scraper API or Browser API, with extraction handled by your workflow or by an agreed delivery scope.

Is a news result the same as a complete article?

No. Discovery is not the same as full article extraction. A result can expose a headline, destination URL, source, snippet, and displayed publication cue without exposing the complete article body. Article fields require a separate eligible public page and source-specific extraction contract.

Which article fields can be delivered?

A scoped article record can include the source URL, canonical URL when exposed, headline, standfirst, byline, displayed publication time, section, visible body, media references, tags, language, and provenance. Field presence varies by publisher and template, so source-agnostic normalized schemas are pilot first.

How are publication time and collection time represented?

published_at records the time attributed to the story by the source when one is exposed. observed_at records when the collection produced the observation. They remain separate because a story may be discovered, revised, or collected long after publication.

Can duplicate coverage of the same event be grouped?

A scheduled or managed pilot can evaluate URL canonicalization, publisher identity, headline similarity, entity cues, and time windows to create candidate story clusters. The original articles remain separate, and ambiguous cluster membership is retained for review rather than silently merged.

Can article updates and revisions be monitored?

A recurring scope can retain successive observations, content hashes, first-seen and last-seen times, and selected field changes. Revision history begins when monitoring begins unless a separate historical source is validated, and source corrections are not interpreted beyond what the public page shows.

How are missing bodies and collection failures handled?

The record can distinguish a search result with no article collection, a field not exposed by the page, a page unavailable in the requested context, a collection failure, a changed extraction pattern, and an excluded source. These states are not collapsed into an empty article body.

Are paywalled or login-only articles included?

Paywalled, login-only, member-only, and otherwise restricted article bodies are not standard scope. A public discovery result or public metadata page does not authorize collection of restricted content behind it.

Can article text be redistributed without restriction?

No blanket copyrighted redistribution right is implied. Customers remain responsible for source terms, copyright, licensing, retention, access controls, and permitted downstream use. A delivery can be limited to metadata, links, excerpts, or other agreed fields when appropriate.

Who maintains source changes?

Proxy customers maintain their own collectors and parsers. WebScrapingAPI maintains access within documented API boundaries. Scheduled feeds and Managed Web Data can include contracted extraction, schema maintenance, quality monitoring, and delivery operations for the agreed source set.

News and articles

Start with discovery access or scope a maintained story feed.

Bring the queries, sources, fields, contexts, and use case. We will help separate what is documented from what the representative pilot needs to prove.