linux

mirror of https://github.com/torvalds/linux.git synced 2024-11-27 14:41:39 +00:00

History

Daniel Rosenberg 2c2eb7a300 f2fs: Support case-insensitive file name lookups Modeled after commit `b886ee3e77` ("ext4: Support case-insensitive file name lookups") """ This patch implements the actual support for case-insensitive file name lookups in f2fs, based on the feature bit and the encoding stored in the superblock. A filesystem that has the casefold feature set is able to configure directories with the +F (F2FS_CASEFOLD_FL) attribute, enabling lookups to succeed in that directory in a case-insensitive fashion, i.e: match a directory entry even if the name used by userspace is not a byte per byte match with the disk name, but is an equivalent case-insensitive version of the Unicode string. This operation is called a case-insensitive file name lookup. The feature is configured as an inode attribute applied to directories and inherited by its children. This attribute can only be enabled on empty directories for filesystems that support the encoding feature, thus preventing collision of file names that only differ by case. * dcache handling: For a +F directory, F2Fs only stores the first equivalent name dentry used in the dcache. This is done to prevent unintentional duplication of dentries in the dcache, while also allowing the VFS code to quickly find the right entry in the cache despite which equivalent string was used in a previous lookup, without having to resort to ->lookup(). d_hash() of casefolded directories is implemented as the hash of the casefolded string, such that we always have a well-known bucket for all the equivalencies of the same string. d_compare() uses the utf8_strncasecmp() infrastructure, which handles the comparison of equivalent, same case, names as well. For now, negative lookups are not inserted in the dcache, since they would need to be invalidated anyway, because we can't trust missing file dentries. This is bad for performance but requires some leveraging of the vfs layer to fix. We can live without that for now, and so does everyone else. * on-disk data: Despite using a specific version of the name as the internal representation within the dcache, the name stored and fetched from the disk is a byte-per-byte match with what the user requested, making this implementation 'name-preserving'. i.e. no actual information is lost when writing to storage. DX is supported by modifying the hashes used in +F directories to make them case/encoding-aware. The new disk hashes are calculated as the hash of the full casefolded string, instead of the string directly. This allows us to efficiently search for file names in the htree without requiring the user to provide an exact name. * Dealing with invalid sequences: By default, when a invalid UTF-8 sequence is identified, ext4 will treat it as an opaque byte sequence, ignoring the encoding and reverting to the old behavior for that unique file. This means that case-insensitive file name lookup will not work only for that file. An optional bit can be set in the superblock telling the filesystem code and userspace tools to enforce the encoding. When that optional bit is set, any attempt to create a file name using an invalid UTF-8 sequence will fail and return an error to userspace. * Normalization algorithm: The UTF-8 algorithms used to compare strings in f2fs is implemented in fs/unicode, and is based on a previous version developed by SGI. It implements the Canonical decomposition (NFD) algorithm described by the Unicode specification 12.1, or higher, combined with the elimination of ignorable code points (NFDi) and full case-folding (CF) as documented in fs/unicode/utf8_norm.c. NFD seems to be the best normalization method for F2FS because: - It has a lower cost than NFC/NFKC (which requires decomposing to NFD as an intermediary step) - It doesn't eliminate important semantic meaning like compatibility decompositions. Although: - This implementation is not completely linguistic accurate, because different languages have conflicting rules, which would require the specialization of the filesystem to a given locale, which brings all sorts of problems for removable media and for users who use more than one language. """ Signed-off-by: Daniel Rosenberg <drosen@google.com> Reviewed-by: Chao Yu <yuchao0@huawei.com> Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>		2019-08-23 07:57:13 -07:00
..
acl.c	f2fs: Replace spaces with tab	2019-05-08 21:23:11 -07:00
acl.h	f2fs: add SPDX license identifiers	2018-09-12 13:07:10 -07:00
checkpoint.c	f2fs: add a rw_sem to cover quota flag changes	2019-07-02 15:40:41 -07:00
data.c	f2fs: support fiemap() for directory inode	2019-08-23 07:57:11 -07:00
debug.c	fs: f2fs: Remove unnecessary checks of SM_I(sbi) in update_general_status()	2019-08-23 07:57:12 -07:00
dir.c	f2fs: Support case-insensitive file name lookups	2019-08-23 07:57:13 -07:00
extent_cache.c	f2fs: introduce f2fs_<level> macros to wrap f2fs_printk()	2019-07-02 15:40:40 -07:00
f2fs.h	f2fs: Support case-insensitive file name lookups	2019-08-23 07:57:13 -07:00
file.c	f2fs: Support case-insensitive file name lookups	2019-08-23 07:57:13 -07:00
gc.c	f2fs: fix to read source block before invalidating it	2019-07-26 17:49:04 -07:00
gc.h	f2fs: add SPDX license identifiers	2018-09-12 13:07:10 -07:00
hash.c	f2fs: Support case-insensitive file name lookups	2019-08-23 07:57:13 -07:00
inline.c	f2fs: Support case-insensitive file name lookups	2019-08-23 07:57:13 -07:00
inode.c	f2fs: Support case-insensitive file name lookups	2019-08-23 07:57:13 -07:00
Kconfig	treewide: Add SPDX license identifier - Makefile/Kconfig	2019-05-21 10:50:46 +02:00
Makefile	License cleanup: add SPDX GPL-2.0 license identifier to files with no license	2017-11-02 11:10:55 +01:00
namei.c	f2fs: Support case-insensitive file name lookups	2019-08-23 07:57:13 -07:00
node.c	f2fs: use generic EFSBADCRC/EFSCORRUPTED	2019-07-02 15:40:41 -07:00
node.h	f2fs: check PageWriteback flag for ordered case	2018-12-26 15:16:56 -08:00
recovery.c	f2fs: use generic EFSBADCRC/EFSCORRUPTED	2019-07-02 15:40:41 -07:00
segment.c	f2fs: fix to avoid discard command leak	2019-08-23 07:57:11 -07:00
segment.h	f2fs: use generic EFSBADCRC/EFSCORRUPTED	2019-07-02 15:40:41 -07:00
shrinker.c	f2fs: fix sbi->extent_list corruption issue	2018-12-26 15:16:54 -08:00
super.c	f2fs: Support case-insensitive file name lookups	2019-08-23 07:57:13 -07:00
sysfs.c	f2fs: include charset encoding information in the superblock	2019-08-23 07:57:13 -07:00
trace.c	f2fs: do not use mutex lock in atomic context	2019-03-05 19:58:06 -08:00
trace.h	f2fs: add SPDX license identifiers	2018-09-12 13:07:10 -07:00
xattr.c	f2fs: fix to detect cp error in f2fs_setxattr()	2019-08-23 07:57:11 -07:00
xattr.h	f2fs: fix to avoid accessing xattr across the boundary	2019-05-09 09:43:29 -07:00