mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* confusion and case problems: utf8 <-> iocharset
@ 2006-07-13  7:56 Eduard Bloch
  2006-07-13 16:25 ` Andrey Borzenkov
  2006-07-13 18:39 ` OGAWA Hirofumi
  0 siblings, 2 replies; 3+ messages in thread
From: Eduard Bloch @ 2006-07-13  7:56 UTC (permalink / raw)
  To: 'Linux Kernel Mailing List'

Hello (to whom it may concern),

I try to understand how the charset mapping with VFAT/Joliet and I found
some inconsistencies between the user expectations, the docs, and the
actuall behaviour.

Users view:

VFAT, NTFS and Joliet use a Unicode charset for storing the names
internaly. That names are mapped to smaller charsets and traditional
encodings like latin1 by the filesystem driver (the iocharset option).
UTF-8 is a Unicode encoding to get the whole charset trough multibyte.

Users expectation:

The way of mapping can be configured with mount options.

The trouble:

First, the terminology in vfat.txt is not consistent with what actually
happens. It says "iocharset" but in fact it is not a charset used for IO
operations, it does not stand for charset at all but for the mapping of
encodings. The better name should be "visible_encoding", IMO.
And in the kernel setup, why do I need a separate "VFAT_IOCHARSET"
option? Why should I not use the systemwide settings, AFAICS that change
is relevant for what the users see and this thing should be consistent
across all mounted filesystems. So why do I need a separate kernel
setting here? Questions over questions.

Second: 
there is the "utf8" option. How does that exactly differ from
iocharset=utf8? There is not clear explanation in vfat.txt. What happens
if you use both options, especially if iocharset!=utf8? Which one is
prefered?

Third:
how can I disable all that funny letter case conversions? They are not
described anywhere properly, nor the way to disable them. IMO there are
two problems:

 - what you write to the FS is not the same what "ls" shows you later.
   Eg. ABW becomes "abw" but "ABWÖ" becomes "ABWÖ". Abcd becomes "Abcd"
   but "ABC" becomes "abc".  Does it make sense? NO.
   I would like to stop the kernel playing such games, I had enough of
   such trouble back in my Windows 98 times.

 - this case conversion can actually break things. When iocharset=utf-8
   and utf8 are used, then you cannot access the data with the same
   name after storing it.

zombie:/tmp# uname -a
Linux zombie 2.6.17.4 #4 Fri Jul 7 12:16:37 CEST 2006 x86_64 GNU/Linux
zombie:/tmp# mount test.img test -o loop,iocharset=utf8,utf8
zombie:/tmp# ls test/test
zombie:/tmp# rm test/test -r
zombie:/tmp# mkdir test/TEST
zombie:/tmp# ls test
test
zombie:/tmp# ls test/test
zombie:/tmp# ls test/TEST
ls: test/TEST: No such file or directory

Full history below.

Thanks,
Eduard.


zombie:/tmp# mkfs.vfat test.img
mkfs.vfat 2.11 (12 Mar 2005)
zombie:/tmp# mount test.img test
mount: test.img is not a block device (maybe try `-o loop'?)
zombie:/tmp# mount test.img test -o loop
zombie:/tmp# grep test /proc/mounts 
/dev/loop/0 /tmp/test vfat rw,fmask=0022,dmask=0022,codepage=cp437,iocharset=iso8859-1 0 0
zombie:/tmp# mkdir test/TEST
zombie:/tmp# ls test/test
zombie:/tmp# ls test/
test
zombie:/tmp# ls test/TEST
zombie:/tmp# umount test.img
zombie:/tmp# mount test.img test -o loop,iocharset=utf-8
mount: wrong fs type, bad option, bad superblock on /dev/loop0,
       missing codepage or other error
       In some cases useful info is found in syslog - try
       dmesg | tail  or so

zombie:/tmp# mount test.img test -o loop,iocharset=utf8
zombie:/tmp# ls test
test
zombie:/tmp# rm test/test -r
zombie:/tmp# mkdir test/TEST
zombie:/tmp# ls test
test
zombie:/tmp# ls test/test
zombie:/tmp# umount test
(reverse-i-search)`u': umount test
zombie:/tmp# mount test.img test -o loop,iocharset=utf8,utf8
zombie:/tmp# ls test/test
zombie:/tmp# rm test/test -r
zombie:/tmp# mkdir test/TEST
zombie:/tmp# ls test
test
zombie:/tmp# ls test/test
zombie:/tmp# ls test/TEST
ls: test/TEST: No such file or directory
zombie:/tmp# umount test
zombie:/tmp# mount test.img test -o loop,iocharset=utf8
zombie:/tmp# ls test/TEST
ls: test/TEST: No such file or directory
zombie:/tmp# ls test/


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: confusion and case problems: utf8 <-> iocharset
  2006-07-13  7:56 confusion and case problems: utf8 <-> iocharset Eduard Bloch
@ 2006-07-13 16:25 ` Andrey Borzenkov
  2006-07-13 18:39 ` OGAWA Hirofumi
  1 sibling, 0 replies; 3+ messages in thread
From: Andrey Borzenkov @ 2006-07-13 16:25 UTC (permalink / raw)
  To: Eduard Bloch, linux-kernel

Eduard Bloch wrote:

> Hello (to whom it may concern),
> 
> I try to understand how the charset mapping with VFAT/Joliet and I found
> some inconsistencies between the user expectations, the docs, and the
> actuall behaviour.
> 
> Users view:
> 
> VFAT, NTFS and Joliet use a Unicode charset for storing the names
> internaly. 

Nope. VFAT is using short name as long as it complies with MSDOS; this name
is stored in codepage character set.

[...]
> 
> Second:
> there is the "utf8" option. How does that exactly differ from
> iocharset=utf8? There is not clear explanation in vfat.txt. What happens
> if you use both options, especially if iocharset!=utf8? Which one is
> prefered?
>

You actually need both; utf8 just says it is OK not to try to mangle names;
it does not substitute iocharset option.
 
> Third:
> how can I disable all that funny letter case conversions? They are not
> described anywhere properly,

man mount

> nor the way to disable them. 

man mount

> IMO there are  
> two problems:
> 
>  - what you write to the FS is not the same what "ls" shows you later.
>    Eg. ABW becomes "abw" but "ABWÖ" becomes "ABWÖ". Abcd becomes "Abcd"
>    but "ABC" becomes "abc".  Does it make sense? NO.

tell this to Microsoft. It is how VFAT works. Although umlaut should
probably be considered as MSDOS-safe, at least with proper codepage option.


>    I would like to stop the kernel playing such games, I had enough of
>    such trouble back in my Windows 98 times.
> 
>  - this case conversion can actually break things. When iocharset=utf-8
>    and utf8 are used, then you cannot access the data with the same
>    name after storing it.
> 

mount shortname=mixed is probably what you want.

> zombie:/tmp# mkdir test/TEST
> zombie:/tmp# ls test
> test
> zombie:/tmp# ls test/test
> zombie:/tmp# ls test/TEST
> ls: test/TEST: No such file or directory

{pts/1}% sudo mount -t vfat -o
loop,shortname=mixed,utf8,uid=bor /var/tmp/test.dos /tmp/x
{pts/1}% LC_ALL=C ll /tmp/x/test
total 2
drwxr-xr-x  2 bor root 2048 Jul 13 20:06 TEST/
{pts/1}% LC_ALL=C ll /tmp/x/test/TEST
total 0
-rwxr-xr-x  1 bor root 0 Jul 13 20:06 ?*
-rwxr-xr-x  1 bor root 0 Jul 13 20:06 ?*
-rwxr-xr-x  1 bor root 0 Jul 13 20:06 ?*
-rwxr-xr-x  1 bor root 0 Jul 13 20:06 ?*

The filesystem is foobared already because it contains both capital and
lowercase versions of the same names. This is what happens when you try to
use wrong codepage.

-andrey


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: confusion and case problems: utf8 <-> iocharset
  2006-07-13  7:56 confusion and case problems: utf8 <-> iocharset Eduard Bloch
  2006-07-13 16:25 ` Andrey Borzenkov
@ 2006-07-13 18:39 ` OGAWA Hirofumi
  1 sibling, 0 replies; 3+ messages in thread
From: OGAWA Hirofumi @ 2006-07-13 18:39 UTC (permalink / raw)
  To: Eduard Bloch; +Cc: 'Linux Kernel Mailing List'

Eduard Bloch <edi@gmx.de> writes:

> The trouble:
>
> First, the terminology in vfat.txt is not consistent with what actually
> happens. It says "iocharset" but in fact it is not a charset used for IO
> operations, it does not stand for charset at all but for the mapping of
> encodings. The better name should be "visible_encoding", IMO.
> And in the kernel setup, why do I need a separate "VFAT_IOCHARSET"
> option? Why should I not use the systemwide settings, AFAICS that change
> is relevant for what the users see and this thing should be consistent
> across all mounted filesystems. So why do I need a separate kernel
> setting here? Questions over questions.

Probably, you want to use "utf8" systemwidely. But, you shouldn't use
utf8 for vfat, because it's breaking. The main reason is this.

> Second: 
> there is the "utf8" option. How does that exactly differ from
> iocharset=utf8? There is not clear explanation in vfat.txt. What happens
> if you use both options, especially if iocharset!=utf8? Which one is
> prefered?

iocharset=utf8 doesn't have a case conversion table.

utf8 option is similar to iocharset=utf8, but utf8 uses case
conversion table of iocharset=xxx. But, there is a known bug.

> Third:
> how can I disable all that funny letter case conversions? They are not
> described anywhere properly, nor the way to disable them. IMO there are
> two problems:
>
>  - what you write to the FS is not the same what "ls" shows you later.
>    Eg. ABW becomes "abw" but "ABWÖ" becomes "ABWÖ". Abcd becomes "Abcd"
>    but "ABC" becomes "abc".  Does it make sense? NO.
>    I would like to stop the kernel playing such games, I had enough of
>    such trouble back in my Windows 98 times.

Probably, you want to use shortname=xxx option.

>  - this case conversion can actually break things. When iocharset=utf-8
>    and utf8 are used, then you cannot access the data with the same
>    name after storing it.

Yes, it's a known bug.
-- 
OGAWA Hirofumi <hirofumi@mail.parknet.co.jp>

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2006-07-13 18:39 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2006-07-13  7:56 confusion and case problems: utf8 <-> iocharset Eduard Bloch
2006-07-13 16:25 ` Andrey Borzenkov
2006-07-13 18:39 ` OGAWA Hirofumi

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®