BlockchainAppMaker

Enterprise Blockchain

IPFS Blockchain Development: Storing Off-Chain Data the Right Way

Blockchains are terrible places to store files, so most dApps store large data off-chain and put a reference on-chain. IPFS makes that reference trustworthy: a content identifier (CID) is derived from the data itself, so anyone can verify that what they fetched is exactly what the contract points to. What IPFS does not do on its own is guarantee the data stays available. That is the job of pinning, and getting it wrong is the most common way IPFS-based projects break.

What IPFS is and is not

The InterPlanetary File System is a set of protocols for addressing, routing and transferring data by content rather than by location. Instead of asking "give me the file at this server's path," you ask the network "who has the data with this hash?" and verify the answer yourself.

  • It is content-addressed, verifiable, deduplicated and fetchable from any node or gateway that has the data.
  • It is not a blockchain, it has no built-in payment or storage guarantee, and it is not private. Anything you add and announce can be fetched by anyone who knows the CID.

That combination makes IPFS a natural partner for blockchains: the chain provides ownership and ordering, IPFS provides tamper-evident storage for the heavy data.

How content addressing works

When you add a file, an IPFS implementation splits it into chunks (256 KiB by default), hashes each chunk, and links them into a Merkle DAG whose root hash becomes the file's CID. Directories are DAG nodes linking to files. The format for files and directories is UnixFS.

A CID packages the hash with metadata about how to interpret it:

PartMeaningExample values
VersionCIDv0 or CIDv1v0 strings start with Qm; v1 in base32 typically starts with bafy or bafk
MultibaseText encoding of the CIDbase58btc (v0), base32 (default for v1)
CodecHow the referenced bytes are structureddag-pb (UnixFS), raw, dag-cbor, dag-json
MultihashHash function and digestsha2-256 by default

An important consequence: the same file can produce different CIDs depending on chunk size, CID version, codec and whether raw leaves are used. If your backend computes a CID one way and a user's tool computes it another way, they will not match even though the bytes are identical. Fix these parameters in your pipeline and document them. Use CIDv1 for new work; it is case-insensitive in base32, which matters for subdomain gateways.

Finding and fetching data

Nodes advertise which CIDs they hold through a distributed hash table (Kademlia-based) and exchange blocks with protocols such as Bitswap, or over HTTP from trustless gateways. In practice most users never run a node; they reach IPFS through an HTTP gateway, using either path style (/ipfs/<cid>/) or subdomain style (<cid>.ipfs.<gateway-domain>). Subdomain gateways give each CID its own browser origin, which matters if you serve HTML or JavaScript.

A regular gateway asks you to trust that it returned the right bytes. A trustless gateway returns raw blocks or CAR files so the client can verify hashes itself, and browser libraries can do that verification for you. If your app's integrity depends on IPFS, verify in the client rather than trusting a gateway.

Pinning: keeping data available

Every IPFS node caches data it fetches and periodically garbage-collects anything not pinned. If the only node that pinned your file goes offline, the CID still exists but nobody can retrieve the content. That is how NFTs end up pointing at nothing.

OptionHow it worksFits
Self-hosted Kubo nodesRun the Go implementation and pin your CIDsTeams with ops capacity wanting full control
IPFS ClusterCoordinates pinsets and replication across several of your own nodesProduction redundancy across regions
Pinning servicesCommercial providers pin for a fee, many through the standard Pinning Service APIMost dApps; quick start, no ops
Filecoin storageStorage providers commit to keep data for a term, with cryptographic proofs on the Filecoin chainLong-term or archival guarantees
Multiple of the abovePin with two independent providers plus your own nodeAnything where loss would be unacceptable

A sound default for a production dApp is to pin with at least two independent parties and keep the original files and CAR exports in your own backup. Monitor retrievability from outside your infrastructure, not just pin status.

Using IPFS with smart contracts

NFT metadata and media

For ERC-721 and ERC-1155 tokens, store the metadata JSON and media on IPFS and return ipfs://<cid>/<tokenId>.json from tokenURI. Use the ipfs:// scheme on-chain rather than a gateway URL, so the reference does not depend on one company's domain; wallets and marketplaces resolve it through their preferred gateway. Upload a whole collection as a directory so a single base CID covers it, and reference media inside the JSON with ipfs:// URIs too. This pattern is covered further in NFT token development.

Reveals and mutability

A CID cannot change, which is the point. If you need a delayed reveal, publish a commitment (the final directory CID or a hash of it) before minting, then switch the base URI once. If metadata must evolve, as with game items, consider IPNS (mutable names signed by a key) or an on-chain base URI controlled by a multisig, and be transparent with holders that the data is mutable.

Documents and records

Enterprise and DeFi apps often store a CID on-chain as a commitment to a document: a loan agreement, an audit report, a governance proposal. The chain proves the document existed in that exact form at that block. Store just the hash or CID; the document itself can live on IPFS or in private storage.

Private and regulated data

IPFS has no access control. If data is sensitive, encrypt it before adding it and manage keys separately, for example with threshold key-management networks or your own key service. Even encrypted, assume the ciphertext is permanent once widely replicated, so do not put personal data on IPFS where erasure rights apply. Keep personal data in conventional storage and anchor only hashes. The EMR and EHR guide discusses this pattern for health records.

Tooling in 2026

  • Kubo: the reference Go implementation (formerly go-ipfs), with an HTTP RPC API.
  • Helia: the JavaScript implementation that replaced js-ipfs, for Node.js and browsers, plus verified-fetch libraries for browser retrieval.
  • IPFS Cluster: pinset orchestration across nodes.
  • CAR files: the archive format for moving DAGs between systems and providers.
  • DNSLink: maps a domain to a CID via a DNS TXT record, useful for hosting dApp front ends.

The IPFS documentation covers each in detail.

Common failure modes

  • Single pinning provider. A provider changes pricing, shuts down or deletes unpaid data, and every token that relied on it breaks at once.
  • Hardcoded gateway URLs. Contracts or metadata pointing at one company's gateway domain fail when that gateway rate-limits or closes.
  • Mismatched CIDs. A reveal or audit fails because the CID computed at upload differs from the one committed on-chain.
  • Serving dApp code from a shared path gateway. Scripts from different CIDs share one browser origin, so one site can read another's local storage.

A build checklist

  1. Decide what goes on-chain (CIDs, hashes, ownership) and what goes on IPFS.
  2. Fix CID parameters (CIDv1, chunker, raw leaves) and compute CIDs locally before uploading, so you never trust a provider's reported CID blindly.
  3. Choose pinning redundancy and set up retrieval monitoring.
  4. Encrypt anything sensitive before it touches IPFS.
  5. Reference content with ipfs:// on-chain and let clients pick gateways.
  6. Host the front end on IPFS with DNSLink if censorship resistance matters to your users, alongside conventional hosting.

Effort and cost

Adding IPFS storage to an existing dApp with a pinning service is days to a couple of weeks of engineering. Running your own redundant cluster with monitoring, or building a product whose core is decentralized storage, is an infrastructure project measured in months and needs ongoing operations. Storage costs depend on volume and provider; for most NFT collections they are small next to the cost of losing the data. For broader dApp architecture, see dApp development and Web3 dApp development.

Frequently asked questions

Is data on IPFS permanent?

Only as long as someone pins it. The CID is permanent, but availability depends on nodes choosing to keep the data. Use redundant pinning or Filecoin deals for long-term storage.

What is the difference between IPFS and Filecoin?

IPFS is the protocol for addressing and moving content. Filecoin is a blockchain-based market where storage providers are paid and must prove they keep data stored. They are commonly used together.

Should my NFT contract store gateway URLs or ipfs:// URIs?

Use ipfs:// URIs. Gateway URLs tie your tokens to one domain that may change or shut down.

Why do I get a different CID for the same file?

CID version, chunk size, codec and raw-leaves settings all change the result. Use the same settings everywhere you compute CIDs.

Can I delete a file from IPFS?

You can unpin it from your nodes, but you cannot force other nodes or caches to delete copies. Treat anything added to IPFS as potentially permanent.

Can a dApp front end be hosted on IPFS?

Yes. Build a static site, add it to IPFS, pin it, and point a domain at it with DNSLink. Use relative paths and hash-based routing so it works under different gateways.