본문으로 건너뛰기

huggingface/transformers.js

State-of-the-art Machine Learning for the web. Run 🤗 Transformers directly in your browser, with no need for a server!

https://skillcdn.ai/gh/huggingface/transformers.js

내 AI에 이 주소를 연결하면 이 스킬을 쓸 수 있어요. 연결 방법 보기

  • 미확인
  • 기본 브랜치main
  • 커밋836b9cb
  • 라이선스Apache-2.0
GitHub에서 보기

transformers-js

Run state-of-the-art machine learning models directly in JavaScript. `@huggingface/transformers` supports text, vision, audio, and multimodal tasks in browsers and Node.js / Bun / Deno, with WebGPU or WASM execution.

경로
.ai/skills/transformers-js/SKILL.md
라이선스
Apache-2.0
호환성
Node.js 20+ (or equivalent Bun / Deno), or a modern browser with ES modules. WebGPU requires runtime and hardware support; WASM is the fallback. Model downloads from the Hugging Face Hub require network access unless you ship models locally.
author
huggingface
repository
https://github.com/huggingface/transformers.js

transformers.js

ML inference for JavaScript, without a Python server. Supports text, vision, audio, and multimodal tasks through a single pipeline() entry point.

Install

npm install @huggingface/transformers

Quick start

import { pipeline } from "@huggingface/transformers";

const classifier = await pipeline("sentiment-analysis");
const output = await classifier("I love transformers!");
// [{ label: "POSITIVE", score: 0.9998 }]

pipeline(task, model?, options?) is the one function you need 90% of the time. Passing no model uses the default for that task.

Supported tasks

<!-- @generated:start id=task-list --> <!-- @generated:end id=task-list -->

For full recipes — every task, grouped by modality, with runnable code — see references/TASKS.md.

Choosing a model

Browse models compatible with transformers.js on the Hub: https://huggingface.co/models?library=transformers.js

Filter by task with the pipeline_tag parameter, e.g. https://huggingface.co/models?library=transformers.js&pipeline_tag=text-generation.

Before recommending a model, confirm it actually has ONNX weights — the library cannot load a model without them. Two ways to check:

  1. Open the model page on the Hub and look for an onnx/ directory in the "Files and versions" tab.
  2. Programmatically, with ModelRegistry.get_available_dtypes(modelId) — returns the list of dtypes shipped. An empty array means the model exists but ships no ONNX files. It throws a ModelFileNotFoundError if the model does not exist or is not accessible (private/gated — the Hub cannot tell these apart), and a regular error on network failures, so an unreachable model is never mistaken for one without ONNX files.
import { ModelRegistry, ModelFileNotFoundError } from "@huggingface/transformers";

// "Xenova/some-model" is a placeholder — substitute the ID you want to check.
try {
  const dtypes = await ModelRegistry.get_available_dtypes("Xenova/some-model");
  if (dtypes.length === 0) {
    // Model exists but has no ONNX files — not usable with transformers.js.
  }
} catch (e) {
  if (e instanceof ModelFileNotFoundError) {
    // Model does not exist, or is private/gated without a token.
  } else {
    throw e; // network failure — retry rather than blacklisting the model
  }
}

Don't suggest a model without verifying this; the failure mode at runtime is a download error that's harder to diagnose than a pre-flight check. For a fuller pre-flight pattern (cache checks, dtype fallback), see references/CONFIGURATION.md.

Quantization

Most pipelines accept a dtype option. Smaller dtypes download and run faster at the cost of some accuracy:

dtypeSizeUse when
fp32LargestMaximum accuracy, Node.js with lots of RAM
fp16~50% of fp32GPU / WebGPU inference
q8~25% of fp32Good default for browsers
q4~12% of fp32Tight memory budgets, large language models
q4f16~12% of fp32Like q4 but with fp16 activations — pairs well with WebGPU LLMs
const pipe = await pipeline("text-generation", "onnx-community/Qwen3-0.6B-ONNX", {
  dtype: "q4",
});
Device

Default is CPU/WASM. Pass device: "webgpu" to run on the GPU when available:

const pipe = await pipeline("sentiment-analysis", null, { device: "webgpu" });

Memory management

Pipelines hold onto model weights and backend sessions. Always call pipe.dispose() when you're done with one — especially in long-running servers, before loading a replacement, or on component unmount.

const pipe = await pipeline("sentiment-analysis");
try {
  const result = await pipe("Great!");
} finally {
  await pipe.dispose();
}

Configuration

The env export lets you control model sources, caching, logging, and the fetch function.

import { env, LogLevel } from "@huggingface/transformers";

env.allowRemoteModels = true;
env.useFSCache = true;           // Node.js: cache downloaded models on disk
env.useBrowserCache = true;      // Browser: cache via the Cache API
env.logLevel = LogLevel.WARNING;

See references/CONFIGURATION.md for the full set of environment options, cache management, and private / gated models.

Pipeline options

Every pipeline accepts a progress_callback for download progress plus options controlling device, dtype, and caching. The per-task recipes in references/TASKS.md show common call options (e.g. top_k, max_new_tokens) in use; shared loading options, generation parameters, streaming, and KV-cache reuse are documented in references/PIPELINE_OPTIONS.md. For the exhaustive per-task option types, see the API reference.

Things to never do

  • Don't reuse a disposed pipeline. Create a new one with pipeline(...) after dispose().
  • Don't recreate pipelines inside hot loops. Create once, call many times.
  • Don't block startup on model downloads. Show progress via progress_callback.
  • Don't fabricate model IDs. Confirm a model exists on the Hub and has ONNX files (look for an onnx/ directory in the repo) before suggesting it to a user.

Reference documentation

This skill's local references:

  • TASKS.md — recipes for every task, grouped by modality (generated)
  • CONFIGURATION.md — env options, caching, model inspection
  • PIPELINE_OPTIONS.md — common pipeline options, dtype, device, generation parameters

이 스킬의 파일

지원 파일 전체 둘러보기