Apify-Js-Sdk - Llms-Txt
Pages: 1113
TaskCollectionListOptions
URL: llms-txt#taskcollectionlistoptions
Contents:
Properties**
****optionaldesc
****optionallimit
****optionaloffset
BuildCollectionClientListOptions
URL: llms-txt#buildcollectionclientlistoptions
Contents:
Properties**
****optionaldesc
****optionallimit
****optionaloffset
Logging
URL: llms-txt#logging
Contents:
The Apify SDK is logging useful information through the logging module from Python's standard library, into the logger with the name apify.
Automatic configuration
When you create an Actor from an Apify-provided template, either in Apify Console or through the Apify CLI, you do not have to configure the logger yourself. The template already contains initialization code for the logger,which sets the logger level to DEBUG and the log formatter to ActorLogFormatter.
Manual configuration
Configuring the log level
In Python's default behavior, if you don't configure the logger otherwise, only logs with level WARNING or higher are printed out to the standard output, without any formatting. To also have logs with DEBUG and INFO level printed out, you need to call the Logger.setLevel method on the logger, with the desired minimum level as an argument.
Configuring the log formatting
By default, only the log message is printed out to the output, without any formatting. To have a nicer output, with the log level printed in color, the messages nicely aligned, and extra log fields printed out,you can use the ActorLogFormatter class from the apify.log module.
Example log configuration
To configure and test the logger, you can use this snippet:
This configuration will cause all levels of messages to be printed to the standard output, with some pretty formatting.
Logger usage
Here you can see how all the log levels would look like.
You can use the extra argument for all log levels, it's not specific to the warning level. When you use Logger.exception, there is no need to pass the Exception object to the log manually, it will automatiacally infer it from the current execution context and print the exception details.
Redirect logs from other Actor runs
In some situations, one Actor is going to start one or more other Actors and wait for them to finish and produce some results. In such cases, you might want to redirect the logs and status messages of the started Actors runs back to the parent Actor run, so that you can see the progress of the started Actors' runs in the parent Actor's logs. This guide will show possibilities on how to do it.
Redirecting logs from Actor.call
Typical use case for log redirection is to call another Actor using the Actor.call method. This method has an optional logger argument, which is by default set to the default literal. This means that the logs of the called Actor will be automatically redirected to the parent Actor's logs with default formatting and filtering. If you set the logger argument to None, then no log redirection happens. The third option is to pass your own Logger instance with the possibility to define your own formatter, filter, and handler. Below you can see those three possible ways of log redirection when starting another Actor run through Actor.call.
Each default redirect logger log entry will have a specific format. After the timestamp, it will contain cyan colored text that will contain the redirect information - the other actor's name and the run ID. The rest of the log message will be printed in the same manner as the parent Actor's logger is configured.
The log redirection can be deep, meaning that if the other actor also starts another actor and is redirecting logs from it, then in the top-level Actor, you can see it as well. See the following example screenshot of the Apify log console when one actor recursively starts itself (there are 2 levels of recursion in the example).
Redirecting logs from already running Actor run
In some cases, you might want to connect to an already running Actor run and redirect its logs to your current Actor run. This can be done using the ApifyClient and getting the streamed log from a specific Actor run. You can then use it as a context manager, and the log redirection will be active in the context, or you can control the log redirection manually by explicitly calling start and stop methods.
You can further decide whether you want to redirect just new logs of the ongoing Actor run, or if you also want to redirect historical logs from that Actor's run, so all logs it has produced since it was started. Both options are shown in the example code below.
Examples:
Example 1 (unknown):
import logging
from apify.log import ActorLogFormatter
async def main() -> None:
handler = logging.StreamHandler()
handler.setFormatter(ActorLogFormatter())
apify_logger = logging.getLogger('apify')
apify_logger.setLevel(logging.DEBUG)
apify_logger.addHandler(handler)
Example 2 (unknown):
import logging
from apify import Actor
from apify.log import ActorLogFormatter
async def main() -> None:
handler = logging.StreamHandler()
handler.setFormatter(ActorLogFormatter())
apify_logger = logging.getLogger('apify')
apify_logger.setLevel(logging.DEBUG)
apify_logger.addHandler(handler)
async with Actor:
Actor.log.debug('This is a debug message')
Actor.log.info('This is an info message')
Actor.log.warning('This is a warning message', extra={'reason': 'Bad Actor!'})
Actor.log.error('This is an error message')
try:
raise RuntimeError('Ouch!')
except RuntimeError:
Actor.log.exception('This is an exceptional message')
Example 3 (unknown):
DEBUG This is a debug message
INFO This is an info message
WARN This is a warning message ({"reason": "Bad Actor!"})
ERROR This is an error message
ERROR This is an exceptional message
Traceback (most recent call last):
File "main.py", line 6, in <module>
raise RuntimeError('Ouch!')
RuntimeError: Ouch!
Example 4 (unknown):
import logging
from apify import Actor
async def main() -> None:
async with Actor:
# Default redirect logger
await Actor.call(actor_id='some_actor_id')
# No redirect logger
await Actor.call(actor_id='some_actor_id', logger=None)
# Custom redirect logger
await Actor.call(
actor_id='some_actor_id', logger=logging.getLogger('custom_logger')
)
Get limits
URL: llms-txt#get-limits
Contents:
- Responses
Returns a complete summary of your account's limits. It is the same information you will see on your account's https://console.apify.com/billing#/limits. The returned data includes the current usage cycle, a summary of your limits, and your current usage.
Examples:
Example 1 (unknown):
GET
https://api.apify.com/v2/users/me/limits
Finding links
URL: llms-txt#finding-links
Contents:
- Extracting links 🔗
- Extracting link URLs in Node.js
- Next Up
Learn what a link looks like in HTML and how to find and extract their URLs when web scraping using both DevTools and Node.js.
Many kinds of links exist on the internet, and we'll cover all the types in the advanced Academy courses. For now, let's think of links as https://developer.mozilla.org/en-US/docs/Web/HTML/Element/a with `` tags. A typical link looks like this:
On a webpage, the link above will look like this: https://example.com When you click it, your browser will navigate to the URL in the `` tag's href attribute (https://example.com).
hrefmeans Hypertext REFerence. You don't need to remember this - just know thathreftypically means some sort of link.
Extracting links 🔗
If a link is an HTML element, and the URL is an attribute, this means that we can extract links the same way as we extracted data. To test this theory in the browser, we can try running the following code in our DevTools console on any website.
Go to the https://warehouse-theme-metal.myshopify.com/collections/sales, open the DevTools Console, paste the above code and run it.
Boom 💥, all the links from the page have now been printed to the console. Most of the links point to other parts of the website, but some links lead to other domains like facebook.com or instagram.com.
Extracting link URLs in Node.js
DevTools Console is a fun playground, but Node.js is way more useful. Let's create a new file in our project called crawler.js and add some basic crawling code that prints all the links from the https://warehouse-theme-metal.myshopify.com/collections/sales.
We'll start from a boilerplate that's very similar to the scraper we built in https://docs.apify.com/academy/web-scraping-for-beginners/data-extraction/node-js-scraper.md.
Aside from importing libraries and downloading HTML, we load the HTML into Cheerio and then use it to retrieve all the `` elements. After that, we iterate over the collected links and print their href attributes, which we access using the https://cheerio.js.org/docs/api/classes/Cheerio#attr method.
When you run the above code, you'll see quite a lot of links in the terminal. Some of them may look wrong, because they don't start with the regular https:// protocol. We'll learn what to do with them in the following lessons.
The https://docs.apify.com/academy/web-scraping-for-beginners/crawling/filtering-links.md will teach you how to select and filter links, so that your crawler will always work only with valid and useful URLs.
Examples:
Example 1 (unknown):
This is a link to example.com
Example 2 (unknown):
// Select all the elements.
const links = document.querySelectorAll('a');
// For each of the links...
for (const link of links) {
// get the value of its 'href' attribute...
const url = link.href;
// and print it to console.
console.log(url);
}
Example 3 (unknown):
import * as cheerio from 'cheerio';
import { gotScraping } from 'got-scraping';
const storeUrl = 'https://warehouse-theme-metal.myshopify.com/collections/sales';
const response = await gotScraping(storeUrl);
const html = response.body;
const $ = cheerio.load(html);
// ------- new code below
const links = $('a');
for (const link of links) {
const url = $(link).attr('href');
console.log(url);
}
DatasetMetadata
URL: llms-txt#datasetmetadata
Contents:
Model for a dataset metadata.
-
- DatasetMetadata
Properties**
****accessed_at
**accessed_at: Annotated[datetime, Field(alias='accessedAt')]
Inherited from StorageMetadata.accessed_at
The timestamp when the storage was last accessed.
****created_at
**created_at: Annotated[datetime, Field(alias='createdAt')]
Inherited from StorageMetadata.created_at
The timestamp when the storage was created.
****id
**id: Annotated[str, Field(alias='id')]
Inherited from StorageMetadata.id
The unique identifier of the storage.
****item_count
The number of items in the dataset.
****model_config
**model_config: Undefined
Overrides StorageMetadata.model_config
****modified_at
**modified_at: Annotated[datetime, Field(alias='modifiedAt')]
Inherited from StorageMetadata.modified_at
The timestamp when the storage was last modified.
****name
**name: Annotated[str | None, Field(alias='name', default=None)]
Inherited from StorageMetadata.name
The name of the storage.
MetamorphOptions
URL: llms-txt#metamorphoptions
Contents:
Properties**
****optionalbuild
Tag or number of the target Actor build to metamorph into (e.g. beta or 1.2.345). If not provided, the run uses build tag or number from the default Actor run configuration (typically latest).
****optionalcontentType
Content type for the input. If not specified, input is expected to be an object that will be stringified to JSON and content type set to application/json; charset=utf-8. If options.contentType is specified, then input must be a String or Buffer.
ActorCollectionClient
URL: llms-txt#actorcollectionclient
Contents:
Sub-client for manipulating Actors.
-
- ActorCollectionClient
Methods**
****__init__
- ****__init__**(*, base_url, root_client, http_client, resource_id, resource_path, params): None
- Overrides ResourceCollectionClient.__init__
Initialize a new instance.
keyword-onlybase_url: str
Base URL of the API server.
keyword-onlyroot_client: ApifyClient
The ApifyClient instance under which this resource client exists.
keyword-onlyhttp_client: HTTPClient
The HTTPClient instance to be used in this client.
optionalkeyword-onlyresource_id: str | None = None
ID of the manipulated resource, in case of a single-resource client.
keyword-onlyresource_path: str
Path to the resource's endpoint on the API server.
optionalkeyword-onlyparams: dict | None = None
Parameters to include in all requests from this client.
****create
- **create(*, name, title, description, seo_title, seo_description, versions, restart_on_error, is_public, is_deprecated, is_anonymously_runnable, categories, default_run_build, default_run_max_items, default_run_memory_mbytes, default_run_timeout_secs, example_run_input_body, example_run_input_content_type, actor_standby_is_enabled, actor_standby_desired_requests_per_actor_run, actor_standby_max_requests_per_actor_run, actor_standby_idle_timeout_secs, actor_standby_build, actor_standby_memory_mbytes): dict
- Create a new Actor.
https://docs.apify.com/api/v2#/reference/actors/actor-collection/create-actor
keyword-onlyname: str
The name of the Actor.
optionalkeyword-onlytitle: str | None = None
The title of the Actor (human-readable).
optionalkeyword-onlydescription: str | None = None
The description for the Actor.
optionalkeyword-onlyseo_title: str | None = None
The title of the Actor optimized for search engines.
optionalkeyword-onlyseo_description: str | None = None
The description of the Actor optimized for search engines.
optionalkeyword-onlyversions: list[dict] | None = None
The list of Actor versions.
optionalkeyword-onlyrestart_on_error: bool | None = None
If true, the Actor run process will be restarted whenever it exits with a non-zero status code.
optionalkeyword-onlyis_public: bool | None = None
Whether the Actor is public.
optionalkeyword-onlyis_deprecated: bool | None = None
Whether the Actor is deprecated.
optionalkeyword-onlyis_anonymously_runnable: bool | None = None
Whether the Actor is anonymously runnable.
optionalkeyword-onlycategories: list[str] | None = None
The categories to which the Actor belongs to.
optionalkeyword-onlydefault_run_build: str | None = None
Tag or number of the build that you want to run by default.
optionalkeyword-onlydefault_run_max_items: int | None = None
Default limit of the number of results that will be returned by runs of this Actor, if the Actor is charged per result.
optionalkeyword-onlydefault_run_memory_mbytes: int | None = None
Default amount of memory allocated for the runs of this Actor, in megabytes.
optionalkeyword-onlydefault_run_timeout_secs: int | None = None
Default timeout for the runs of this Actor in seconds.
optionalkeyword-onlyexample_run_input_body: Any = None
Input to be prefilled as default input to new users of this Actor.
optionalkeyword-onlyexample_run_input_content_type: str | None = None
The content type of the example run input.
optionalkeyword-onlyactor_standby_is_enabled: bool | None = None
Whether the Actor Standby is enabled.
optionalkeyword-onlyactor_standby_desired_requests_per_actor_run: int | None = None
The desired number of concurrent HTTP requests for a single Actor Standby run.
optionalkeyword-onlyactor_standby_max_requests_per_actor_run: int | None = None
The maximum number of concurrent HTTP requests for a single Actor Standby run.
optionalkeyword-onlyactor_standby_idle_timeout_secs: int | None = None
If the Actor run does not receive any requests for this time, it will be shut down.
optionalkeyword-onlyactor_standby_build: str | None = None
The build tag or number to run when the Actor is in Standby mode.
optionalkeyword-onlyactor_standby_memory_mbytes: int | None = None
The memory in megabytes to use when the Actor is in Standby mode.
****list
- **list(*, my, limit, offset, desc, sort_by): ListPage[dict]
- List the Actors the user has created or used.
https://docs.apify.com/api/v2#/reference/actors/actor-collection/get-list-of-actors
optionalkeyword-onlymy: bool | None = None
If True, will return only Actors which the user has created themselves.
optionalkeyword-onlylimit: int | None = None
How many Actors to list.
optionalkeyword-onlyoffset: int | None = None
What Actor to include as first when retrieving the list.
optionalkeyword-onlydesc: bool | None = None
Whether to sort the Actors in descending order based on their creation date.
optionalkeyword-onlysort_by: Literal[createdAt, stats.lastRunStartedAt] | None = 'createdAt'
Field to sort the results by.
Returns ListPage[dict]
Properties**
****http_client
**http_client: HTTPClient | HTTPClientAsync
Inherited from BaseClient.http_client
Overrides _BaseBaseClient.http_client
****params
Inherited from _BaseBaseClient.params
****resource_id
**resource_id: str | None
Inherited from _BaseBaseClient.resource_id
****root_client
**root_client: ApifyClient | ApifyClientAsync
Inherited from BaseClient.root_client
Overrides _BaseBaseClient.root_client
****url
Inherited from _BaseBaseClient.url
Delete request
URL: llms-txt#delete-request
Contents:
- Request
- Responses
Clientshttps://docs.apify.com/api/client/js/reference/class/RequestQueueClient#deleteDeletes given request from queue.
Examples:
Example 1 (unknown):
DELETE
https://api.apify.com/v2/request-queues/:queueId/requests/:requestId
ActorEnvVarClient
URL: llms-txt#actorenvvarclient
Contents:
Sub-client for manipulating a single Actor environment variable.
-
- ActorEnvVarClient
Methods**
****__init__
- ****__init__**(*, base_url, root_client, http_client, resource_id, resource_path, params): None
- Overrides ResourceClient.__init__
Initialize a new instance.
keyword-onlybase_url: str
Base URL of the API server.
keyword-onlyroot_client: ApifyClient
The ApifyClient instance under which this resource client exists.
keyword-onlyhttp_client: HTTPClient
The HTTPClient instance to be used in this client.
optionalkeyword-onlyresource_id: str | None = None
ID of the manipulated resource, in case of a single-resource client.
keyword-onlyresource_path: str
Path to the resource's endpoint on the API server.
optionalkeyword-onlyparams: dict | None = None
Parameters to include in all requests from this client.
****delete
- **delete(): None
- Delete the Actor environment variable.
****get
- **get(): dict | None
- Return information about the Actor environment variable.
https://docs.apify.com/api/v2#/reference/actors/environment-variable-object/get-environment-variable
Returns dict | None
****update
- **update(*, is_secret, name, value): dict
- Update the Actor environment variable with specified fields.
optionalkeyword-onlyis_secret: bool | None = None
Whether the environment variable is secret or not.
keyword-onlyname: str
The name of the environment variable.
keyword-onlyvalue: str
The value of the environment variable.
Properties**
****http_client
**http_client: HTTPClient | HTTPClientAsync
Inherited from BaseClient.http_client
Overrides _BaseBaseClient.http_client
****params
Inherited from _BaseBaseClient.params
****resource_id
**resource_id: str | None
Inherited from _BaseBaseClient.resource_id
****root_client
**root_client: ApifyClient | ApifyClientAsync
Inherited from BaseClient.root_client
Overrides _BaseBaseClient.root_client
****url
Inherited from _BaseBaseClient.url
RequestQueueClientUnlockRequestsResult
URL: llms-txt#requestqueueclientunlockrequestsresult
Contents:
Properties**
****unlockedCount
**unlockedCount: number
Pricing and costs
URL: llms-txt#pricing-and-costs
Contents:
- Computing your costs for PPE and PPR Actors
- Discount tiers and pricing strategy
- Implementing discount tiers
- Additional benefits and enterprise tiers
Learn how to set Actor pricing and calculate your costs, including platform usage rates, discount tiers, and profit formulas for PPE and PPR monetization models.
Computing your costs for PPE and PPR Actors
For both PPE and PPR Actors, profit is computed using the formula (0.8 * revenue) - costs. In this section, we'll explain how the costs component is calculated.
When paying users run your Actor, it generates platform usage in the form of compute units, data traffic, API operati
…(truncated)