From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ed1-f51.google.com (mail-ed1-f51.google.com [209.85.208.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 19E8D3655D8 for ; Mon, 27 Jul 2026 06:18:15 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.51 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785133098; cv=none; b=KqL/VVNOMIHI0C6rL97yxBuLL6681iVDRPHRmKPdKcP2Olj4ekw2uuohR1WDhxcTbHO0/R5lIYERz68lB1kOTks3du2/9szRxmumAMok2tJftckOYMRbXwXgf7BQGSjfz9PARGXzK+Mn4AMiQGlpSo1w6Jz7nYLhajDg1tqKIig= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785133098; c=relaxed/simple; bh=4RnSXRlFTtcvpuBj2UrW/EbdX0woNfnxDksjMbfaxjI=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=cLw6OcE0ANoo8znTvWUxuezof4crTEuiAXMlfjTDDxY7D74vgULiMgpZ/a+x21q0ji/0rcs8Tddepxe64pnT5IisT02qq/XAXY1lwR95MEX9IYH9nP54ECGNQamzJ0Bcz5TzpfgOPF610kUCGABQ9gZ1kXiryR3dDFQCz5jCzcU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com; spf=pass smtp.mailfrom=suse.com; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b=Mge754mV; arc=none smtp.client-ip=209.85.208.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=suse.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=suse.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=suse.com header.i=@suse.com header.b="Mge754mV" Received: by mail-ed1-f51.google.com with SMTP id 4fb4d7f45d1cf-69af1bc780bso197321a12.1 for ; Sun, 26 Jul 2026 23:18:15 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=suse.com; s=google; t=1785133094; x=1785737894; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=kv1usJVqUmBl0SRtdM1BtqDPTmR0aDfg3c/KtIYHeqI=; b=Mge754mVDk5/oZpgidab0bHUJBnxpRRfXpadjOzAs4ZOtNYUYy2vbhoLUrT2V0rR5A i1ggA/2CMGigu3DBWkH97aQxPMMWhFyqYkUcjNCJE1Ywe2xPwO1aCbFlYg/vyUQnaXHK CHNL7K2i78khkLas0iZg2BG5UgiCy4gefQOte4mhd2FYdJTYqZ9uiTifc7HBNyAwM29r QQlvNfbEFPw+mo06MduV8Gpfjt4FcfUi+zvh8VvOlx4BC+zvzaAn0SW3lxyT1z8oY39d jtV0HEIturLGD9Bu7pj5gaoqPQ8kNAvFP49IQWUgNsgUSx4rqrpgR/7qj3RAxP1L+yG+ 31JA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785133094; x=1785737894; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=kv1usJVqUmBl0SRtdM1BtqDPTmR0aDfg3c/KtIYHeqI=; b=B18tX5o/KoTzYHY74dC0+JoD5ftUO2bPjfiDV+nw8wAj6vJ3W0QQ1CbQS7utuJRTQv rDYMmAAR4krFXuoA5Lh/esD7nekcZCQsTIkwEbTXavZ7AYyRgqf2lkirVNY+NKJNBQEQ u/oH58ZSyeUyAWTIUsQk82QglYIxty0m70mwdll5PLH0aAP6V6lsz6lL8vn+LnQ+oXub Hg4DH6osn/B+4I2BjCDNOUSI5qoclXDGhPKWO/CHQAgMgno/jPwVjRpGS9u4D6f9Wr7J bXCCBs/V/JGlDs93xiVTPEaWxaD3j3wmh+GPuA6tFM6QRnGG0qYwFtcKjvnyjM050k2V dXdg== X-Forwarded-Encrypted: i=1; AHgh+RoAAYGnnpDZwro5/8XirpoJZcNvp9F8iG+hP1leN7LfbNRqrPnGWioAO7CCmf9FEWYKoWMjTFSPRsrISQ0=@vger.kernel.org X-Gm-Message-State: AOJu0YxyHNzdK3DtZlIjA2d92QM8B9UNNjVg0Vdw1mPi4fmkr0mws38X bf39Hs3gJ2J9JWJn1Bj3I9Mn/ObFIauU4OIPHpgU3VJJxX++TYMkgwnV8qS/+eriglo= X-Gm-Gg: AR+sD11xxqiMfHsJWHDKqVGipNUOX3zke92bJLIPppvLG3ApynzhGIqxUz2geRAKjL1 6aR+YCwlWH0UaUaTtBBcQ1Tbwi/qudCD6MOuI/nDES0Q5Hl1MzKSzYY7aN31gZOtGLKu0el1yTy Hu6juu58tBbKBWy/MLHrflKKczc6lq9oJ1lcCcyvr2Uvc+9mZP3vJL98FVSMyUDIylsHStFQhjz cHwqlKlKWSjLgnV1tkAOyhCe58WcT6MCBu7VZqxp/83gm16fW3zcSoY0cFnkOBDhITbXJaN/v4R ZAUJIA5Rlw33rFpTMJJ8Df07TnJuGhJvAwj9ejJSo6nfz0wQRg0YnmAhwk7GpFB+dnXiYeXI8Nn F4dJuy3UO7ZKDbyaLQ6ezktHNJdMsNyfgt9gfu52UTzDjUZpFgZEBbHVh/yePejm5JpkVBHzveQ FD7bs= X-Received: by 2002:a05:6402:501a:b0:698:c152:f69d with SMTP id 4fb4d7f45d1cf-69fc1176a22mr1717704a12.7.1785133094288; Sun, 26 Jul 2026 23:18:14 -0700 (PDT) Received: from p15.suse.cz ([202.127.77.110]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-38f2951f0afsm2583300a91.16.2026.07.26.23.18.11 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sun, 26 Jul 2026 23:18:13 -0700 (PDT) From: Heming Zhao To: joseph.qi@linux.alibaba.com, mark@fasheh.com, jlbec@evilplan.org, hch@lst.de Cc: Heming Zhao , ocfs2-devel@lists.linux.dev, linux-kernel@vger.kernel.org Subject: [RFC PATCH v2 1/4] ocfs2: Add new ocfs2_map_blocks() to introduce iomap feature Date: Mon, 27 Jul 2026 14:17:57 +0800 Message-ID: <20260727061802.18485-2-heming.zhao@suse.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260727061802.18485-1-heming.zhao@suse.com> References: <20260727061802.18485-1-heming.zhao@suse.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit As part of migrating OCFS2 DIO read/write code paths towards the modern and high-performant iomap framework, this patch introduces an iomap API to replace old high-overhead VFS buffer_head structure paths. This patch establishes the foundational block mapping routines required by subsequent iomap integration patches. The implementation draws inspiration from ext4_map_blocks(). Signed-off-by: Heming Zhao --- fs/ocfs2/aops.c | 94 +++++++++++++++++++++++++++++++++++++++ fs/ocfs2/aops.h | 2 + fs/ocfs2/buffer_head_io.c | 12 ----- fs/ocfs2/ocfs2.h | 45 ++++++++++++++++++- 4 files changed, 140 insertions(+), 13 deletions(-) diff --git a/fs/ocfs2/aops.c b/fs/ocfs2/aops.c index 4acdbb70882c..08df5e3b5196 100644 --- a/fs/ocfs2/aops.c +++ b/fs/ocfs2/aops.c @@ -126,6 +126,100 @@ static int ocfs2_lock_get_block(struct inode *inode, sector_t iblock, return ret; } +int ocfs2_map_blocks(struct inode *inode, struct ocfs2_map_block *map, + int flags) +{ + int err = 0; + unsigned int ext_flags; + u64 max_blocks = map->len; + u64 p_blkno, count, past_eof; + struct ocfs2_super *osb = OCFS2_SB(inode->i_sb); + int create = flags & OCFS2_GET_BLOCKS_CREATE; + + if (OCFS2_I(inode)->ip_flags & OCFS2_INODE_SYSTEM_FILE) + mlog(ML_NOTICE, "map_block on system inode 0x%p (%llu)\n", + inode, inode->i_ino); + + if (S_ISLNK(inode->i_mode)) { + /* + * TODO: refer ocfs2_get_block() to handle + * ocfs2_read_folio in the future + */ + mlog(ML_NOTICE, "map_block on S_ISLNK file, node 0x%p (%llu)\n", + inode, inode->i_ino); + dump_stack(); + goto bail; + } + + err = ocfs2_extent_map_get_blocks(inode, map->lblk, &p_blkno, &count, + &ext_flags); + if (err) { + mlog(ML_ERROR, "get_blocks() failed, inode: 0x%p, " + "block: %llu\n", inode, map->lblk); + goto bail; + } + + if (max_blocks < count) + count = max_blocks; + + map->pblk = p_blkno; + map->len = count; + + /* + * ocfs2 never allocates in this function - the only time we + * need to use MAP_NEW is when we're extending i_size on a file + * system which doesn't support holes, in which case MAP_NEW + * allows __block_write_begin() to zero. + * + * If we see this on a sparse file system, then a truncate has + * raced us and removed the cluster. In this case, we clear + * the buffers dirty and uptodate bits and let the buffer code + * ignore it as a hole. + */ + if (create && map->pblk == 0 && ocfs2_sparse_alloc(osb)) { + map->flags &= ~(OCFS2_MAP_DIRTY | OCFS2_MAP_UPTODATE); + goto bail; + } + + if (p_blkno) { + if (ext_flags & OCFS2_EXT_UNWRITTEN) { + map->flags |= OCFS2_MAP_UNWRITTEN; + } else if (!(ext_flags & OCFS2_EXT_UNWRITTEN)) { + /* Treat the unwritten extent as a hole for zeroing purposes. */ + map->flags |= OCFS2_MAP_MAPPED; + } else { + /* nothing to do */ + } + } + + if (!ocfs2_sparse_alloc(osb)) { + if (map->pblk == 0) { + err = -EIO; + mlog(ML_ERROR, + "iblock = %llu p_blkno = %llu blkno=(%llu)\n", + (unsigned long long)map->lblk, + (unsigned long long)map->pblk, + (unsigned long long)OCFS2_I(inode)->ip_blkno); + mlog(ML_ERROR, "Size %llu, clusters %u\n", + (unsigned long long)i_size_read(inode), + OCFS2_I(inode)->ip_clusters); + dump_stack(); + goto bail; + } + } + + past_eof = ocfs2_blocks_for_bytes(inode->i_sb, i_size_read(inode)); + + if (create && (map->lblk >= past_eof)) + map->flags |= OCFS2_MAP_NEW; + +bail: + if (err < 0) + return -EIO; + else + return map->len; +} + int ocfs2_get_block(struct inode *inode, sector_t iblock, struct buffer_head *bh_result, int create) { diff --git a/fs/ocfs2/aops.h b/fs/ocfs2/aops.h index 114efc9111e4..8dd6edd7c1a1 100644 --- a/fs/ocfs2/aops.h +++ b/fs/ocfs2/aops.h @@ -42,6 +42,8 @@ int ocfs2_size_fits_inline_data(struct buffer_head *di_bh, u64 new_size); int ocfs2_get_block(struct inode *inode, sector_t iblock, struct buffer_head *bh_result, int create); +int ocfs2_map_blocks(struct inode *inode, struct ocfs2_map_block *map, + int flags); /* all ocfs2_dio_end_io()'s fault */ #define ocfs2_iocb_is_rw_locked(iocb) \ test_bit(0, (unsigned long *)&iocb->private) diff --git a/fs/ocfs2/buffer_head_io.c b/fs/ocfs2/buffer_head_io.c index 7bfe377af2df..493f2209cca5 100644 --- a/fs/ocfs2/buffer_head_io.c +++ b/fs/ocfs2/buffer_head_io.c @@ -23,18 +23,6 @@ #include "buffer_head_io.h" #include "ocfs2_trace.h" -/* - * Bits on bh->b_state used by ocfs2. - * - * These MUST be after the JBD2 bits. Hence, we use BH_JBDPrivateStart. - */ -enum ocfs2_state_bits { - BH_NeedsValidate = BH_JBDPrivateStart, -}; - -/* Expand the magic b_state functions */ -BUFFER_FNS(NeedsValidate, needs_validate); - int ocfs2_write_block(struct ocfs2_super *osb, struct buffer_head *bh, struct ocfs2_caching_info *ci) { diff --git a/fs/ocfs2/ocfs2.h b/fs/ocfs2/ocfs2.h index 62cad6522c7a..095f7ae5dded 100644 --- a/fs/ocfs2/ocfs2.h +++ b/fs/ocfs2/ocfs2.h @@ -509,7 +509,50 @@ struct ocfs2_super struct ocfs2_filecheck_sysfs_entry osb_fc_ent; }; -#define OCFS2_SB(sb) ((struct ocfs2_super *)(sb)->s_fs_info) +/* + * Bits on bh->b_state used by ocfs2. + * + * These MUST be after the JBD2 bits. Hence, we use BH_JBDPrivateStart. + */ +enum ocfs2_state_bits { + BH_NeedsValidate = BH_JBDPrivateStart, +}; + +/* Expand the magic b_state functions */ +BUFFER_FNS(NeedsValidate, needs_validate); + +/* + * Logical to physical block mapping, used by ocfs2_map_blocks() + * + * This structure is used to pass requests into ocfs2_map_blocks() as + * well as to store the information returned by ocfs2_map_blocks(). It + * takes less room on the stack than a struct buffer_head. + */ +#define OCFS2_MAP_NEW BIT(BH_New) +#define OCFS2_MAP_MAPPED BIT(BH_Mapped) +#define OCFS2_MAP_UNWRITTEN BIT(BH_Unwritten) +/* useless? #define OCFS2_MAP_BOUNDARY BIT(BH_Boundary) */ +/* useless? #define OCFS2_MAP_DELAYED BIT(BH_Delay) */ +#define OCFS2_MAP_DIRTY BIT(BH_Dirty) +#define OCFS2_MAP_UPTODATE BIT(BH_Uptodate) +#define OCFS2_MAP_NEEDS_VALIDATE BIT(BH_NeedsValidate) +#define OCFS2_MAP_DEFER_COMPLETION BIT(BH_Defer_Completion) +#define OCFS2_MAP_FLAGS (OCFS2_MAP_NEW | OCFS2_MAP_MAPPED |\ + OCFS2_MAP_DIRTY | OCFS2_MAP_UPTODATE |\ + OCFS2_MAP_NEEDS_VALIDATE |\ + OCFS2_MAP_DEFER_COMPLETION) + +struct ocfs2_map_block { + u64 pblk; /* physical block# */ + u64 lblk; /* logical block# */ + u64 len; /* number of block */ + unsigned int flags; +}; + +/* Flags used by ocfs2_map_blocks() */ +#define OCFS2_GET_BLOCKS_CREATE (0x0001) + +#define OCFS2_SB(sb) ((struct ocfs2_super *)(sb)->s_fs_info) /* Useful typedef for passing around journal access functions */ typedef int (*ocfs2_journal_access_func)(handle_t *handle, -- 2.54.0